Yi Ru Wang1 · Carter Ung1 · Jiafei Duan1,2 · Markus Grotz1 · Wilbert Pumacay1 · Ranjay Krishna1,2 · Dieter Fox1 · Siddhartha Srinivasa1
1University of Washington · 2Allen Institute for AI
Binary success rates hide how a bimanual policy fails. This challenge asks you to build policies that are not just successful but well-behaved — coordinated, efficient, safe, and robust to spatial variation — and scores them with stage-level diagnostics rather than a single pass/fail bit.
Part of the CoRL 2026 Embodied AI Benchmark and Sim-to-Real Challenge Suite, hosted at the workshop From GPU-Accelerated Simulation to Scalable and Generalizable Real Robot Policy Learning — November 12, 2026, Austin, Texas, USA.
| Track | Track 2 of 3 — Diagnostic Bimanual Manipulation |
|---|---|
| Simulator | MuJoCo, bimanual Franka Panda |
| Tasks | 8 task families — 28 variations spanning position, rotation, and combined spatial shifts |
| Demonstrations | 3,000+ human demonstrations collected via VR and keyboard teleoperation |
| Scored on | Outcome, Efficiency, Bimanual Coordination, Safety & Stability — 14 metrics total |
| What you submit | Policy checkpoint, evaluation logs, and a short technical report |
| Baselines | ACT, Diffusion Policy, GR00T (see the leaderboard) |
| License | MIT — benchmark, assets, and demonstrations are free to use |
Dates subject to change
Launch and submission dates are being finalized with the workshop organizers. Watch the RoboEval repository or the workshop site for announcements.
Prize amounts, compute credits, and travel grants for this track are being finalized with our sponsors. Award categories are listed below; winners present their approach at the workshop on November 12, 2026.
Open-source bonus
Entries that release training code and checkpoints under a permissive license are eligible for an additional open-source award and are highlighted on the public leaderboard. We strongly encourage reproducible submissions — the point of a diagnostic benchmark is that others can build on what you learned.
In our study, policies with nearly identical success rates diverged sharply in how they executed the same task — some lost alignment, others lost temporally consistent bimanual control. Behavioral metrics correlated with success in over half of all task–metric pairs, and stayed informative even where binary success saturated. This challenge is built around that finding.
Every task decomposes into skill-specific stages, so a failure is localized to where it happened — not just recorded as a zero.
Each task family ships position, rotation, and combined variants that probe spatial generalization rather than memorization of a fixed scene.
Arm velocity mismatch and gripper height offset are first-class metrics, so genuinely bimanual behavior is rewarded over two arms acting independently.
Collisions, slips, and jerk are measured throughout. A policy that succeeds by flailing through the scene will not win.
The challenge uses the RoboEval task suite: everyday bimanual activities drawn from service, warehouse, and industrial settings, each paired with expert demonstrations and stage annotations. Every task family is evaluated across its variations.
| Task Family | Variations | # Demos | Traj. Len | Skills | Coordination Type |
|---|---|---|---|---|---|
| Lift Tray | Static, Pos, Rot, PR | 543 | 67.6 | grasp, lift | Tight Sym. |
| Stack Two Cubes | Static, Pos, Rot, PR | 492 | 107.0 | grasp, hold, place | Loosely Coord. |
| Stack Single Book Shelf | Static, Pos, PR | 202 | 172.3 | push, grasp, lift, place | Loosely Coord. |
| Cube Handover | Static, Pos, Rot, PR | 408 | 93.5 | grasp, hold | Loosely Coord. |
| Lift Pot | Static, Pos, Rot, PR | 176 | 53.2 | grasp, lift | Tight Sym. |
| Pack Box | Static, Pos, Rot, PR | 394 | 133.2 | push | Uncoord. |
| Pick Book From Table | Static, Pos, Rot, PR | 366 | 106.0 | grasp, lift | Loosely Coord. |
| Rotate Valve | Static, Pos, PR | 349 | 119.6 | grasp, rotate along axis | Uncoord. |
Pos = position-only variation · Rot = orientation-only variation · PR = combined position and rotation. Trajectory length is the mean demonstration length in seconds.
Over 3,000 human demonstrations are provided, collected through VR teleoperation on an Oculus Quest and through keyboard teleoperation. Each demonstration includes RGB and depth observations, point clouds, joint positions, object poses, and stage annotations. You may use all of it, a subset, or none of it.
Policies interact through the standard RoboEval interface. Action modes cover joint position (absolute or delta) and end-effector Cartesian control, at a configurable control frequency. Observation modes range from full (RGB, depth, point clouds) to lightweight (joint positions, object poses).
Submissions are ranked on four metric families. Each is reported per task, per variation, and aggregated across the suite. Full definitions, units, and aggregation rules are in the Metrics Reference.
git clone --recurse-submodules git@github.com:Robo-Eval/RoboEval.git
cd RoboEval
conda create -n roboeval python=3.10
conda activate roboeval
pip install -e ".[examples]"
python examples/1_data_replay.py
from roboeval.envs.lift_pot import LiftPotPositionAndOrientation
from roboeval.action_modes import JointPositionActionMode
from roboeval.robots.configs.panda import BimanualPanda
env = LiftPotPositionAndOrientation(
action_mode=JointPositionActionMode(absolute=True, ee=False),
render_mode="human",
control_frequency=20,
robot_cls=BimanualPanda,
)
python examples/5_gather_metrics.py
results.json, and technical report, then
submit the link. Full instructions are on the
Submission Guidelines page.
Documentation
Hardware note
RoboEval runs on MuJoCo and evaluates on CPU or a single consumer GPU. VR teleoperation for collecting your own demonstrations additionally requires an Oculus Quest and GLIBC 2.32+.
Questions?
Open an issue on the RoboEval repository or email the track organizers — see Contact below.
This track is organized by the RoboEval team at the University of Washington and the Allen Institute for AI.
Yi Ru Wang1 · Carter Ung1 · Jiafei Duan1,2 · Markus Grotz1 · Wilbert Pumacay1 · Ranjay Krishna1,2 · Dieter Fox1 · Siddhartha Srinivasa1
1University of Washington · 2Allen Institute for AI
Track questions: yiruwang@cs.washington.edu
Workshop questions: zhenzhenl@nvidia.com
Technical issues: github.com/Robo-Eval/RoboEval/issues
The challenge suite runs three independent tracks under a shared reporting standard. Entering one does not require entering the others.
1,000 household activities across 50 scenes and 10,000 objects in Isaac Sim, with long-horizon navigation and bimanual manipulation.
Stage-level failure diagnosis for bimanual manipulation in MuJoCo, scored on behavior as well as outcome.
Policies trained in GPU-accelerated simulation must generalize to physical robots across multiple real-world embodiments.
If you use RoboEval or participate in this challenge, please cite:
@misc{wang2025roboevalroboticmanipulationmeets,
title={RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation},
author={Yi Ru Wang and Carter Ung and Grant Tannert and Jiafei Duan and Josephine Li
and Amy Le and Rishabh Oswal and Markus Grotz and Wilbert Pumacay
and Yuquan Deng and Ranjay Krishna and Dieter Fox and Siddhartha Srinivasa},
year={2025},
eprint={2507.00435},
archivePrefix={arXiv},
primaryClass={cs.RO}
}