How submissions to the RoboEval Diagnostic Bimanual Challenge are run, measured, and ranked — and what participants are and are not permitted to do.
Protocol status
The structure below is final. Exact numeric settings marked TBC (episode counts, evaluation seeds, and scoring weights) are confirmed when the challenge launches, and the launch announcement is authoritative if the two ever disagree.
Every submission is evaluated on all 8 task families across all 28 variations. There is no task selection: a policy that is excellent on two families and absent on six will rank poorly. Each task–variation pair is run for a fixed number of independent episodes with fixed evaluation seeds.
| Task families | 8 (all mandatory) |
|---|---|
| Variations | 28 (all mandatory) |
| Episodes per variation | TBC — identical for every submission |
| Evaluation seeds | TBC — published as a fixed list at launch |
| Episode termination | On task success, on the per-task step budget, or on an unrecoverable simulator state |
| Control frequency | 20 Hz default; submissions may use a different rate if declared in the report |
| Simulator | MuJoCo, bimanual Franka Panda, as pinned in the challenge release tag |
Evaluation seeds control initial object poses within each variation's distribution. They are published, not secret — the challenge tests spatial generalization, not seed guessing. Overfitting to the published evaluation seeds is nonetheless a rules violation (see Rules), and organizers re-run leading submissions on a held-out seed set to confirm results hold.
Held-out verification
Top-ranked entries are re-evaluated by the organizers on an unpublished seed set drawn from the same variation distributions. A submission whose held-out performance departs substantially from its reported numbers is investigated before any award is confirmed.
Each task decomposes into ordered, skill-specific stages — for example Stack Single Book Shelf runs push → grasp → lift → place. The evaluator records the furthest stage reached in every episode, which produces Task Progression and localizes failures. An episode that grasps correctly but never places is scored differently from one that never made contact, even though both have a success of 0.
Fourteen metrics across four families. ↑ means higher is better, ↓ means lower is better. All are computed by the RoboEval evaluation harness and reported per task, per variation, as mean and standard deviation over episodes.
| Metric | Key | Dir. | Unit | Definition |
|---|---|---|---|---|
| Outcome | ||||
| Success | success | ↑ | fraction | Whether the episode reached the task goal. Binary per episode, averaged over episodes. |
| Task Progression | tp | ↑ | fraction | Furthest stage reached, normalized by the number of stages in the task. Partial credit for partial execution. |
| Efficiency | ||||
| Cartesian Path Length | cpl | ↓ | m | Total 3D distance travelled by the end-effectors over the episode. |
| Joint Path Length | jpl | ↓ | rad | Total angular distance moved, summed across all joints of both arms. |
| Orientation Path Length | opl | ↓ | rad | Integrated end-effector rotation change over the episode. |
| Completion Time | ct | ↓ | s | Wall-clock task time until termination, in simulated seconds. |
| Trajectory Length | tl | ↓ | steps | Number of control steps executed before termination. |
| Bimanual Coordination | ||||
| Bimanual Arm Velocity Difference | bavd | ↓ | m/s | Mean absolute difference in end-effector speed between the two arms. Low values indicate the arms move as a coordinated pair rather than sequentially. |
| Bimanual Gripper Vertical Difference | bgvd | ↓ | m | Mean height offset between the two grippers. Matters most for tightly symmetric tasks such as Lift Tray and Lift Pot, where a tilted carry drops the payload. |
| Safety & Stability | ||||
| Self Collision Count | scc | ↓ | count | Number of contact events between the robot's own links, including arm-to-arm contacts. |
| Environment Collision Count | ecc | ↓ | count | Number of unintended contact events between the robot and scene geometry. |
| Slip Count | sc | ↓ | count | Number of times a grasped object slips or is dropped unintentionally. |
| Mean Joint Jerk | mjj | ↓ | rad/s³ | Average rate of change of joint acceleration — smoothness of the commanded motion. |
| Mean Cartesian Jerk | mcj | ↓ | m/s³ | Average rate of change of end-effector acceleration in 3D. |
Behavioral metrics are conditioned on progress
A policy that does nothing has excellent path length and zero collisions. To prevent this, efficiency, coordination, and safety metrics are computed over the portion of the episode in which the policy is making progress, and the composite score multiplies behavioral quality by outcome. Standing still cannot win.
Raw metric values are not comparable across tasks — Stack Single Book Shelf trajectories are three times longer than Lift Pot trajectories. Each behavioral metric is therefore normalized against the human demonstration distribution for the same task and variation. A policy whose behavior matches expert statistics scores 1.0 on that metric; one that is markedly worse than the expert trends toward 0. Outcome metrics are already in [0, 1] and are used directly.
The leaderboard's primary ranking is the RoboEval Diagnostic Score (RDS): the weighted mean of the four family scores, averaged over all 28 variations with equal weight per variation.
| Outcome | 50% — Success 35%, Task Progression 15% (weights TBC) |
|---|---|
| Safety & Stability | 20% (TBC) |
| Bimanual Coordination | 15% (TBC) |
| Efficiency | 15% (TBC) |
Ties on RDS are broken in order by: (1) mean Success, (2) mean Task Progression, (3) Safety & Stability family score, (4) earlier submission timestamp.
Alongside the aggregate ranking, the site publishes per-family and per-task leaderboards. These do not determine the main prizes, but the Best Coordination award is decided from the coordination family score, and per-task boards are the intended reference for papers reporting on a subset of RoboEval.
results.json by hand.Disqualification
Submissions that violate these rules are removed from the leaderboard. Where a violation appears accidental, organizers contact the team first and allow a corrected resubmission before the deadline. Deliberate circumvention of the evaluation protocol is disqualifying for the full challenge.
Each team may submit up to TBC entries. The most recent valid submission is the one ranked; earlier entries remain visible in the team's history. There is no limit on running the evaluation harness locally.
Every submission must be reproducible by the organizers from the artifacts provided. Concretely:
Full artifact requirements are specified on the Submission Format page.
RoboEval — the benchmark, assets, and demonstration data — is released under the MIT License. You retain full ownership of your policy, code, and checkpoints. Submitting grants the organizers permission to run your submission for evaluation and to publish the resulting metrics, your team name, and your method description on the leaderboard and in any workshop report summarizing the challenge.
Releasing your code and checkpoints is encouraged and qualifies for the open-source award, but is not required to compete.
All participants are expected to follow the CoRL 2026 code of conduct. Report concerns to the track organizers at yiruwang@cs.washington.edu.