Documentation

Evaluation & Rules

How submissions to the RoboEval Diagnostic Bimanual Challenge are run, measured, and ranked — and what participants are and are not permitted to do.

Protocol status

The structure below is final. Exact numeric settings marked TBC (episode counts, evaluation seeds, and scoring weights) are confirmed when the challenge launches, and the launch announcement is authoritative if the two ever disagree.

What gets run

Every submission is evaluated on all 8 task families across all 28 variations. There is no task selection: a policy that is excellent on two families and absent on six will rank poorly. Each task–variation pair is run for a fixed number of independent episodes with fixed evaluation seeds.

Task families 8 (all mandatory)
Variations 28 (all mandatory)
Episodes per variation TBC — identical for every submission
Evaluation seeds TBC — published as a fixed list at launch
Episode termination On task success, on the per-task step budget, or on an unrecoverable simulator state
Control frequency 20 Hz default; submissions may use a different rate if declared in the report
Simulator MuJoCo, bimanual Franka Panda, as pinned in the challenge release tag

Seeds and determinism

Evaluation seeds control initial object poses within each variation's distribution. They are published, not secret — the challenge tests spatial generalization, not seed guessing. Overfitting to the published evaluation seeds is nonetheless a rules violation (see Rules), and organizers re-run leading submissions on a held-out seed set to confirm results hold.

Held-out verification

Top-ranked entries are re-evaluated by the organizers on an unpublished seed set drawn from the same variation distributions. A submission whose held-out performance departs substantially from its reported numbers is investigated before any award is confirmed.

Stage-level scoring

Each task decomposes into ordered, skill-specific stages — for example Stack Single Book Shelf runs push → grasp → lift → place. The evaluator records the furthest stage reached in every episode, which produces Task Progression and localizes failures. An episode that grasps correctly but never places is scored differently from one that never made contact, even though both have a success of 0.

Fourteen metrics across four families. ↑ means higher is better, ↓ means lower is better. All are computed by the RoboEval evaluation harness and reported per task, per variation, as mean and standard deviation over episodes.

Metric Key Dir. Unit Definition
Outcome
Successsuccess↑fraction Whether the episode reached the task goal. Binary per episode, averaged over episodes.
Task Progressiontp↑fraction Furthest stage reached, normalized by the number of stages in the task. Partial credit for partial execution.
Efficiency
Cartesian Path Lengthcpl↓m Total 3D distance travelled by the end-effectors over the episode.
Joint Path Lengthjpl↓rad Total angular distance moved, summed across all joints of both arms.
Orientation Path Lengthopl↓rad Integrated end-effector rotation change over the episode.
Completion Timect↓s Wall-clock task time until termination, in simulated seconds.
Trajectory Lengthtl↓steps Number of control steps executed before termination.
Bimanual Coordination
Bimanual Arm Velocity Differencebavd↓m/s Mean absolute difference in end-effector speed between the two arms. Low values indicate the arms move as a coordinated pair rather than sequentially.
Bimanual Gripper Vertical Differencebgvd↓m Mean height offset between the two grippers. Matters most for tightly symmetric tasks such as Lift Tray and Lift Pot, where a tilted carry drops the payload.
Safety & Stability
Self Collision Countscc↓count Number of contact events between the robot's own links, including arm-to-arm contacts.
Environment Collision Countecc↓count Number of unintended contact events between the robot and scene geometry.
Slip Countsc↓count Number of times a grasped object slips or is dropped unintentionally.
Mean Joint Jerkmjj↓rad/s³ Average rate of change of joint acceleration — smoothness of the commanded motion.
Mean Cartesian Jerkmcj↓m/s³ Average rate of change of end-effector acceleration in 3D.

Behavioral metrics are conditioned on progress

A policy that does nothing has excellent path length and zero collisions. To prevent this, efficiency, coordination, and safety metrics are computed over the portion of the episode in which the policy is making progress, and the composite score multiplies behavioral quality by outcome. Standing still cannot win.

Normalization

Raw metric values are not comparable across tasks — Stack Single Book Shelf trajectories are three times longer than Lift Pot trajectories. Each behavioral metric is therefore normalized against the human demonstration distribution for the same task and variation. A policy whose behavior matches expert statistics scores 1.0 on that metric; one that is markedly worse than the expert trends toward 0. Outcome metrics are already in [0, 1] and are used directly.

Composite score

The leaderboard's primary ranking is the RoboEval Diagnostic Score (RDS): the weighted mean of the four family scores, averaged over all 28 variations with equal weight per variation.

Outcome50% — Success 35%, Task Progression 15%  (weights TBC)
Safety & Stability20%  (TBC)
Bimanual Coordination15%  (TBC)
Efficiency15%  (TBC)

Tie-breaking

Ties on RDS are broken in order by: (1) mean Success, (2) mean Task Progression, (3) Safety & Stability family score, (4) earlier submission timestamp.

Secondary leaderboards

Alongside the aggregate ranking, the site publishes per-family and per-task leaderboards. These do not determine the main prizes, but the Best Coordination award is decided from the coordination family score, and per-task boards are the intended reference for papers reporting on a subset of RoboEval.

Who can enter

  • Open to anyone: academic, industrial, and independent participants are all welcome.
  • Teams may be of any size. Each team submits under a single identity.
  • An individual may belong to only one team.
  • Challenge organizers and their immediate research groups may submit entries for reference, but are not eligible for prizes and are marked as such on the leaderboard.

What is allowed

  • Any policy architecture, including vision-language-action models, diffusion policies, and classical or hybrid controllers.
  • Any pretraining, including on external robotics datasets and internet-scale corpora.
  • Collecting additional demonstrations yourself with the provided teleoperation tools.
  • Data augmentation, synthetic data generation, and domain randomization during training.
  • Task-conditioned or per-task models, provided the same submission covers all 28 variations.
  • Test-time compute such as sampling and reranking, as long as the declared control frequency is met.

What is not allowed

  • Reading privileged simulator state that is not exposed through the standard observation interface — ground-truth object poses in an image-only submission, contact flags, or internal solver state.
  • Modifying the environment, task definition, reward, or termination conditions. The challenge release tag is the reference; local edits to it invalidate a submission.
  • Tuning on the published evaluation seeds. Use them to sanity-check, not to select checkpoints or hyperparameters.
  • Hard-coding trajectories keyed to specific initial states rather than learning or computing a policy from observations.
  • Editing the metrics computation or post-processing results.json by hand.
  • Multiple accounts used to circumvent the submission limit.

Disqualification

Submissions that violate these rules are removed from the leaderboard. Where a violation appears accidental, organizers contact the team first and allow a corrected resubmission before the deadline. Deliberate circumvention of the evaluation protocol is disqualifying for the full challenge.

Submission limits

Each team may submit up to TBC entries. The most recent valid submission is the one ranked; earlier entries remain visible in the team's history. There is no limit on running the evaluation harness locally.

Every submission must be reproducible by the organizers from the artifacts provided. Concretely:

  • The policy checkpoint must load and run against the pinned challenge release with the declared dependencies.
  • The technical report must state the training data used, the architecture, the action and observation modes, the control frequency, and any test-time compute.
  • Randomness in the policy must be seedable, and the seed must be declared.
  • Total training compute should be reported in GPU-hours. This does not affect ranking; it is collected so the community can compare methods fairly.

Full artifact requirements are specified on the Submission Format page.

RoboEval — the benchmark, assets, and demonstration data — is released under the MIT License. You retain full ownership of your policy, code, and checkpoints. Submitting grants the organizers permission to run your submission for evaluation and to publish the resulting metrics, your team name, and your method description on the leaderboard and in any workshop report summarizing the challenge.

Releasing your code and checkpoints is encouraged and qualifies for the open-source award, but is not required to compete.

All participants are expected to follow the CoRL 2026 code of conduct. Report concerns to the track organizers at yiruwang@cs.washington.edu.