CoRL 2026 · Challenge Suite · Track 2

RoboEval Diagnostic Bimanual Challenge

Binary success rates hide how a bimanual policy fails. This challenge asks you to build policies that are not just successful but well-behaved — coordinated, efficient, safe, and robust to spatial variation — and scores them with stage-level diagnostics rather than a single pass/fail bit.

Part of the CoRL 2026 Embodied AI Benchmark and Sim-to-Real Challenge Suite, hosted at the workshop From GPU-Accelerated Simulation to Scalable and Generalizable Real Robot Policy Learning — November 12, 2026, Austin, Texas, USA.

Track Track 2 of 3 — Diagnostic Bimanual Manipulation
Simulator MuJoCo, bimanual Franka Panda
Tasks 8 task families — 28 variations spanning position, rotation, and combined spatial shifts
Demonstrations 3,000+ human demonstrations collected via VR and keyboard teleoperation
Scored on Outcome, Efficiency, Bimanual Coordination, Safety & Stability — 14 metrics total
What you submit Policy checkpoint, evaluation logs, and a short technical report
Baselines ACT, Diffusion Policy, GR00T (see the leaderboard)
License MIT — benchmark, assets, and demonstrations are free to use
  • TBA Challenge launch — benchmark, demonstrations, and baselines released
  • TBA Leaderboard opens for submissions
  • TBA Submission deadline — checkpoints, logs, and reports due
  • November 9, 2026 Winners announced
  • November 12, 2026 Workshop presentation — CoRL 2026, Austin, Texas

Dates subject to change

Launch and submission dates are being finalized with the workshop organizers. Watch the RoboEval repository or the workshop site for announcements.

Prize amounts, compute credits, and travel grants for this track are being finalized with our sponsors. Award categories are listed below; winners present their approach at the workshop on November 12, 2026.

🥇 First Place
TBA
Best overall diagnostic score across all task families
🥈 Second Place
TBA
Runner-up on the aggregate leaderboard
🥉 Third Place
TBA
Third on the aggregate leaderboard
⭐ Best Coordination
TBA
Strongest bimanual coordination and safety profile, independent of rank

Open-source bonus

Entries that release training code and checkpoints under a permissive license are eligible for an additional open-source award and are highlighted on the public leaderboard. We strongly encourage reproducible submissions — the point of a diagnostic benchmark is that others can build on what you learned.

In our study, policies with nearly identical success rates diverged sharply in how they executed the same task — some lost alignment, others lost temporally consistent bimanual control. Behavioral metrics correlated with success in over half of all task–metric pairs, and stayed informative even where binary success saturated. This challenge is built around that finding.

Stage-level diagnosis

Every task decomposes into skill-specific stages, so a failure is localized to where it happened — not just recorded as a zero.

Systematic variation

Each task family ships position, rotation, and combined variants that probe spatial generalization rather than memorization of a fixed scene.

Coordination is scored

Arm velocity mismatch and gripper height offset are first-class metrics, so genuinely bimanual behavior is rewarded over two arms acting independently.

Safety counts

Collisions, slips, and jerk are measured throughout. A policy that succeeds by flailing through the scene will not win.

The challenge uses the RoboEval task suite: everyday bimanual activities drawn from service, warehouse, and industrial settings, each paired with expert demonstrations and stage annotations. Every task family is evaluated across its variations.

Task Family Variations # Demos Traj. Len Skills Coordination Type
Lift TrayStatic, Pos, Rot, PR54367.6 grasp, liftTight Sym.
Stack Two CubesStatic, Pos, Rot, PR492107.0 grasp, hold, placeLoosely Coord.
Stack Single Book ShelfStatic, Pos, PR202172.3 push, grasp, lift, placeLoosely Coord.
Cube HandoverStatic, Pos, Rot, PR40893.5 grasp, holdLoosely Coord.
Lift PotStatic, Pos, Rot, PR17653.2 grasp, liftTight Sym.
Pack BoxStatic, Pos, Rot, PR394133.2 pushUncoord.
Pick Book From TableStatic, Pos, Rot, PR366106.0 grasp, liftLoosely Coord.
Rotate ValveStatic, Pos, PR349119.6 grasp, rotate along axisUncoord.

Pos = position-only variation · Rot = orientation-only variation · PR = combined position and rotation. Trajectory length is the mean demonstration length in seconds.

Demonstration data

Over 3,000 human demonstrations are provided, collected through VR teleoperation on an Oculus Quest and through keyboard teleoperation. Each demonstration includes RGB and depth observations, point clouds, joint positions, object poses, and stage annotations. You may use all of it, a subset, or none of it.

Observation & action interface

Policies interact through the standard RoboEval interface. Action modes cover joint position (absolute or delta) and end-effector Cartesian control, at a configurable control frequency. Observation modes range from full (RGB, depth, point clouds) to lightweight (joint positions, object poses).

Submissions are ranked on four metric families. Each is reported per task, per variation, and aggregated across the suite. Full definitions, units, and aggregation rules are in the Metrics Reference.

Outcome

  • Success (SR)
  • Task Progression (TP)

Efficiency

  • Cartesian Path Length (CPL)
  • Joint Path Length (JPL)
  • Orientation Path Length (OPL)
  • Completion Time (CT)
  • Trajectory Length (TL)

Coordination

  • Bimanual Arm Velocity Difference (BAVD)
  • Bimanual Gripper Vertical Difference (BGVD)

Safety & Stability

  • Self Collision Count (SCC)
  • Environment Collision Count (ECC)
  • Slip Count (SC)
  • Mean Joint Jerk (MJJ)
  • Mean Cartesian Jerk (MCJ)
  1. Install RoboEval Clone the repository with submodules and install into a clean Python 3.10 environment.
    git clone --recurse-submodules git@github.com:Robo-Eval/RoboEval.git
    cd RoboEval
    
    conda create -n roboeval python=3.10
    conda activate roboeval
    
    pip install -e ".[examples]"
  2. Verify the installation Replay a bundled demonstration to confirm the simulator and assets load correctly.
    python examples/1_data_replay.py
  3. Explore the tasks and data Instantiate an environment and inspect the observation and action interface.
    from roboeval.envs.lift_pot import LiftPotPositionAndOrientation
    from roboeval.action_modes import JointPositionActionMode
    from roboeval.robots.configs.panda import BimanualPanda
    
    env = LiftPotPositionAndOrientation(
        action_mode=JointPositionActionMode(absolute=True, ee=False),
        render_mode="human",
        control_frequency=20,
        robot_cls=BimanualPanda,
    )
  4. Train your policy Use the provided demonstrations, your own data, or any pretraining you like. See the rules for what is and is not permitted.
  5. Run the evaluation harness Produce the metrics file that your submission is scored from.
    python examples/5_gather_metrics.py
  6. Submit Package your checkpoint, results.json, and technical report, then submit the link. Full instructions are on the Submission Guidelines page.

Documentation

  1. Evaluation protocol — episodes, seeds, budgets
  2. Metrics reference — all 14 metrics defined
  3. Ranking & aggregation
  4. Rules & eligibility
  5. Submission format
  6. Verification process
  7. FAQ

Hardware note

RoboEval runs on MuJoCo and evaluates on CPU or a single consumer GPU. VR teleoperation for collecting your own demonstrations additionally requires an Oculus Quest and GLIBC 2.32+.

Questions?

Open an issue on the RoboEval repository or email the track organizers — see Contact below.

This track is organized by the RoboEval team at the University of Washington and the Allen Institute for AI.

Yi Ru Wang1 · Carter Ung1 · Jiafei Duan1,2 · Markus Grotz1 · Wilbert Pumacay1 · Ranjay Krishna1,2 · Dieter Fox1 · Siddhartha Srinivasa1

1University of Washington  ·  2Allen Institute for AI

Contact

Track questions: yiruwang@cs.washington.edu
Workshop questions: zhenzhenl@nvidia.com
Technical issues: github.com/Robo-Eval/RoboEval/issues

The challenge suite runs three independent tracks under a shared reporting standard. Entering one does not require entering the others.

Track 1

BEHAVIOR Long-Horizon Household

1,000 household activities across 50 scenes and 10,000 objects in Isaac Sim, with long-horizon navigation and bimanual manipulation.

behavior.stanford.edu ↗

Track 2 — This page

RoboEval Diagnostic Bimanual

Stage-level failure diagnosis for bimanual manipulation in MuJoCo, scored on behavior as well as outcome.

robo-eval.github.io

Track 3

RoboChess Sim-to-Real

Policies trained in GPU-accelerated simulation must generalize to physical robots across multiple real-world embodiments.

RoboChess challenge ↗

If you use RoboEval or participate in this challenge, please cite:

@misc{wang2025roboevalroboticmanipulationmeets,
      title={RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation},
      author={Yi Ru Wang and Carter Ung and Grant Tannert and Jiafei Duan and Josephine Li
              and Amy Le and Rishabh Oswal and Markus Grotz and Wilbert Pumacay
              and Yuquan Deng and Ranjay Krishna and Dieter Fox and Siddhartha Srinivasa},
      year={2025},
      eprint={2507.00435},
      archivePrefix={arXiv},
      primaryClass={cs.RO}
}