Coverage
Every modality. Every embodiment. Every environment.
One verification standard across every stream we capture or label.
Modalities
Every modality.
- RGB · stereo · depth
- video
- Synchronized across cameras, with offsets measured per episode.
- LiDAR · radar
- point cloud
- Captured alongside video for 3D and sensor-fusion labels.
- IMU · proprioception
- time series
- Joint states, actions and inertial data. IMU within ±20 ms of video.
- Force · torque · tactile
- contact
- For manipulation where vision isn’t enough.
- Audio · language
- multimodal
- Stage-level language labels, and audio where the task needs it.
Embodiments
Every embodiment.
- Single and bimanual arms
- manipulation
- Including our own ALOHA bimanual setups.
- Humanoids
- whole-body
- Whole-body demonstrations and teleoperation.
- Mobile manipulators · AMRs
- navigation
- Navigation and manipulation on the move.
- Quadrupeds · drones
- locomotion
- Legged and aerial platforms.
- Human demonstrators
- egocentric
- Head-mounted and handheld rigs, including UMI grippers.
Our own setups are ALOHA bimanual stations. Other embodiments run on your hardware.
Environments
Collected where the work happens.
-
Food & hospitality
Bimanual manipulation, tool use
-
Logistics
Pick, pack, palletize
-
Manufacturing
Assembly, inspection
-
Textile
Deformable objects
-
Retail
Restocking, scanning
-
Workshop
Power tools, measurement
-
Agriculture
Sorting, harvesting
-
Home
Laundry, tidying, cleaning
- Manufacturing · assembly
- industrial
- Assembly benches, inspection stations and production lines.
- Warehouse · logistics
- operations
- Picking, packing and palletizing.
- Retail · hospitality · food
- service
- Shelves, counters and kitchens.
- Healthcare · laboratory
- regulated
- Controlled environments with stricter handling rules.
- Home · agriculture · construction
- field
- Homes, farms and sites outside the lab.
Not on the list?
New sensor, robot or environment?
We scope it the same way: a protocol first, then a graded evaluation batch.