Robotics & embodied AI / Video evaluation

What happened?
Where did it go wrong?

Turn robot attempts and human demonstrations into reviewed evidence. We evaluate visible task outcomes, compare attempts, and correct video annotations so your team can use the results for model evaluation and dataset improvement.

What we evaluate

More than a successful-looking clip.

A demonstration can look plausible while missing a required step or ending in the wrong state. We organize review around the instruction, observable evidence, and your definition of success.

01

Task completion & failure review

Assess whether the recorded attempt follows the instruction and reaches the required end state. Mark visible failures, incomplete steps, and recovery. Record the first observable failure when it can be located, and flag outcomes the footage cannot establish.

02

Attempt comparisons

Compare two attempts against the same task and rubric. Record which better meets the criteria, why, and whether the evidence supports a tie or an uncertain judgment. Preserve differences in starting conditions that affect the comparison.

03

Annotation correction

Review human- or machine-generated descriptions against the footage. Correct action boundaries, timestamps, hand and object references, and claims about task progress. Retain a record of changes so your team can trace each accepted label to the source.

Execution & quality assurance

A defined rubric. A traceable review.

Bring the video, task instructions, and any existing labels. We scope the reviewer skills, annotation format, evidence requirements, and acceptance process with your team.

  1. 01

    Scope the evidence

    Agree on task coverage, camera views, timestamp conventions, label definitions, and which outcomes are observable. Identify any additional context or sensor evidence needed for the requested judgments.

  2. 02

    Calibrate reviewers

    Select reviewers for the task and align them on a sample of clear, difficult, and ambiguous episodes. Specialized procedures require the relevant domain knowledge and customer guidance.

  3. 03

    Review & resolve

    Review the agreed episodes, correct annotations, and document judgments. Route disagreements and uncertain cases through the agreed review process instead of forcing unsupported labels.

  4. 04

    Accept & deliver

    Check schema, timestamps, source references, and agreed quality criteria. Deliver accepted records with review status, remaining exceptions, and version information.

What you receive

Labels, judgments, and the evidence behind them.

Receive a structured dataset in an agreed format, linked to your original footage. Choose the review outputs your model or data workflow needs.

Episode and segment records
Stable video identifiers, task instructions, segment start and end times, action descriptions, and relevant hand or object references.
Outcomes and comparisons
Task-completion judgments, visible failure and recovery labels, pairwise preferences where requested, reasons, and uncertainty or unobservable-outcome flags.
Review and correction history
Accepted annotations, original-to-corrected changes where supplied, rubric version, review status, disagreement resolutions, and documented coverage gaps.

Illustrative annotation correction

Contact is not a successful grasp.

Illustrative example: the task is to lift a cup and place it on a tray. The review uses the visible sequence to distinguish an attempted grasp from a completed action.

01

Original annotation

“The gripper grasps the cup.” The label covers an interval in which the gripper touches the cup, slips, and tries again.

02

Reviewed sequence

Separate the first contact, unsuccessful grasp, second attempt, visible lift, and placement into the appropriate intervals. Record timestamps from the supplied footage.

03

Useful correction

The accepted record distinguishes failure from recovery and identifies the observed end state. Your team can use it to audit labels, evaluate behavior, or select episodes for further work.

Review the evidence your footage contains.

This service evaluates recorded behavior and corrects its annotations. Robot control, teleoperation, new demonstration capture, and sensor collection are separately scoped work. Video review cannot establish hidden forces, internal robot state, or outcomes outside the available views.

Start with a sample and a task

What should this attempt have achieved?

Describe the task, the footage you have, the labels or judgments you need, and the decision the dataset will support. We will scope a sample review, the rubric, and a defined delivery.

Keep the first inquiry high level. Agree on an appropriate transfer and access arrangement before sharing confidential footage or personal information.