Solutions · Evaluation
Human judgement for policy rollouts.
Send us your rollout videos. We return success and failure verdicts, failure causes and blind A/B comparisons, graded to your rubric with agreement reported per batch.
How it works
You run the policy. We judge the result.
Your team runs the policy and records the rollouts. We grade the videos: no hardware to ship, no access to your robot needed.
-
01 · YOU SEND
Rollouts and rubric
Rollout videos, your rubric or success criteria, and which policy produced each rollout, hidden from reviewers for A/B.
-
02 · WE GRADE
Two reviewers
Two independent reviewers grade every rollout.
-
03 · YOU RECEIVE
Verdicts and a report
Verdicts, failure causes and a batch report.
Grading
What we grade
-
Verdicts
Success, failure or retry for every rollout, against criteria written down before grading starts.
-
Failure causes
A taxonomy of why rollouts fail, built with you and applied the same way in every batch.
-
Blind A/B
Reviewers compare rollouts from different policies without knowing which policy produced which.
-
Benchmark runs
Recorded runs of your benchmark tasks, graded to your rubric.
Consistency
Consistent from batch to batch
Rubric first
Criteria and edge cases agreed before the first rollout is graded.
Two reviewers, agreement reported
Every rollout graded independently by two reviewers. Agreement reported per batch, with disagreements listed.
What you receive
- Verdicts and failure causes for every rollout, as JSON
- A batch report with results by task and condition, reviewer agreement and reason codes
Evaluation batch D
Start with a free evaluation batch.
500 policy rollouts graded, in 10 days. Free, and graded before you commit.