Solutions · Evaluation

Human judgement for policy rollouts.

Send us your rollout videos. We return success and failure verdicts, failure causes and blind A/B comparisons, graded to your rubric with agreement reported per batch.

How it works

You run the policy. We judge the result.

Your team runs the policy and records the rollouts. We grade the videos: no hardware to ship, no access to your robot needed.

  1. 01 · YOU SEND

    Rollouts and rubric

    Rollout videos, your rubric or success criteria, and which policy produced each rollout, hidden from reviewers for A/B.

  2. 02 · WE GRADE

    Two reviewers

    Two independent reviewers grade every rollout.

  3. 03 · YOU RECEIVE

    Verdicts and a report

    Verdicts, failure causes and a batch report.

Grading

What we grade

  • Verdicts

    Success, failure or retry for every rollout, against criteria written down before grading starts.

  • Failure causes

    A taxonomy of why rollouts fail, built with you and applied the same way in every batch.

  • Blind A/B

    Reviewers compare rollouts from different policies without knowing which policy produced which.

  • Benchmark runs

    Recorded runs of your benchmark tasks, graded to your rubric.

Consistency

Consistent from batch to batch

Rubric first

Criteria and edge cases agreed before the first rollout is graded.

Two reviewers, agreement reported

Every rollout graded independently by two reviewers. Agreement reported per batch, with disagreements listed.

What you receive

  • Verdicts and failure causes for every rollout, as JSON
  • A batch report with results by task and condition, reviewer agreement and reason codes
Abstract wave of glowing blue particles

Evaluation batch D

Start with a free evaluation batch.

500 policy rollouts graded, in 10 days. Free, and graded before you commit.