Three side-by-side views of a robot arm holding a spatula over omelettes with different browning and folds.
AI-generated illustration: three omelette outcomes show how appearance and folding can vary within the same task.

What decision would the data support?

Start with a model-development decision: can a model represent or predict the actions and state changes that lead from raw ingredients to an accepted omelette, including what happens when the process goes wrong?

That decision determines what must be observed. A collection designed only to recognize a finished omelette would look different from one intended to learn action sequencing, predict future states, or recover from failure.

What is one data unit?

One unit is a complete episode, not an isolated image or label. It connects the starting state, actions, resulting states, outcome, and quality status.

FieldIllustrative example
Initial stateEggs, pan, fat, utensils, heat source, and optional fillings are visible and identified.
ActionCrack, whisk, heat, pour, stir, set, add filling, fold, and plate.
Resulting stateThe observable state after each action, including texture, temperature cues, position, and tool state.
OutcomeAccepted omelette, acceptable variation, recoverable failure, or failed episode.
Quality statusComplete, incomplete, rejected, retake required, or accepted after review.
ProvenanceParticipant, consent, capture conditions, source files, transformations, and review history.

How would the pilot be orchestrated?

  1. Decide

    Specify whether the model must learn action order, future-state prediction, failure recognition, recovery, or outcome evaluation.

  2. Specify

    Define the episode, required views or sensor records, action vocabulary, state annotations, permitted variation, and acceptance test.

  3. Qualify

    Select people who can perform the task consistently and follow the capture protocol without erasing natural variation.

  4. Produce

    Collect successful demonstrations together with planned variations, failures, recoveries, and counterfactual choices.

  5. Control

    Check completeness, synchronization, rights, protocol compliance, annotation consistency, and reviewer disagreement. Reject or retake episodes that fail the acceptance rules.

  6. Deliver

    Package accepted episodes with structured metadata, provenance, quality status, and documentation of the final acceptance test.

Why do variation and failure matter?

A useful dataset would not contain only polished demonstrations. Eggs stick. A shell fragment falls into the bowl. The pan is too cool. A fold tears. A cook changes tools, lowers the heat, or starts again. Those branches connect actions to consequences and show which recoveries still lead to an acceptable outcome.

Counterfactuals can make the same episode more useful: what likely changes if the pan is hotter, the fold begins earlier, or a different tool is used? The point is not to collect every possible omelette. It is to collect the distinctions required by the model decision.

What makes an episode acceptable?

Acceptance is defined before collection. An episode might require a complete action sequence, usable observation quality, synchronized timestamps, permitted ingredients and tools, an outcome judgment, complete provenance, and agreement or adjudication on ambiguous states.

The output is an accepted, client-specific dataset rather than a folder of unreviewed recordings or a count of labor hours.

What would be delivered?

  • Accepted episode media or sensor records.
  • Action, state, outcome, and quality metadata.
  • Failure and recovery classifications.
  • Rights and provenance records.
  • Protocol and acceptance documentation.

A physical task is a sequence of consequences

What does your model need to learn or prove?

Have a physical task or predicted future that your model needs to learn or prove? Contact Consequence Labs.

Scope a pilot