Documentation

Submission Guidelines

You run the RoboEval evaluation harness locally, package the checkpoint and metrics it produces, and submit a link. Organizers verify leading entries by re-running them.

Before you start

Read the rules first. The most common cause of an invalidated submission is an edited environment or metrics file, which is easy to avoid and painful to discover after the deadline.

  1. Pin the challenge release Check out the tagged challenge release so your numbers are comparable with everyone else's.
    # TODO(organizers): replace with the real challenge tag at launch
    git fetch --tags
    git checkout corl2026-challenge
    pip install -e ".[examples]"
  2. Run the evaluation harness Evaluate your policy across all 8 task families and 28 variations using the published evaluation seeds. This writes the per-task, per-variation metrics that your ranking is computed from.
    python examples/5_gather_metrics.py \
        --policy path/to/your_policy.py \
        --checkpoint path/to/checkpoint.pt \
        --output submission/results.json

    Exact flags follow the challenge release; --help is authoritative.

  3. Assemble the submission bundle Collect the checkpoint, metrics, rollout logs, and report into the layout described in Submission Format.
  4. Upload the bundle Host it somewhere the organizers can download without an account request — Google Drive, Hugging Face Hub, S3, or an institutional server. Hugging Face is preferred for checkpoints.
  5. Submit the link through the form One form entry per submission. See Submitting.

Your uploaded bundle must have this structure:

submission/
├── results.json            # REQUIRED — harness output, see schema below
├── report.pdf              # REQUIRED — technical report, max 4 pages + refs
├── metadata.yaml           # REQUIRED — team and method declaration
├── checkpoint/             # REQUIRED — policy weights, loadable as declared
│   └── ...
├── policy/                 # REQUIRED — inference code needed to load and run
│   ├── policy.py           #            the checkpoint against RoboEval
│   └── requirements.txt
└── logs/                   # REQUIRED — raw per-episode evaluation logs
    └── <Task>_<Variation>/
        └── episode_<i>.json

metadata.yaml

team_name: "Your Team"
method_name: "YourMethod-v2"
contact_email: "you@example.edu"
affiliation: "Your Institution"
members:
  - "First Author"
  - "Second Author"

roboeval_release: "corl2026-challenge"
action_mode: "JointPosition(absolute=true, ee=false)"
observation_mode: "full"          # full | lightweight
control_frequency: 20             # Hz
policy_seed: 0

training_data:
  roboeval_demos: true            # did you train on the provided demonstrations?
  additional_demos: 0             # count of self-collected demonstrations
  external_pretraining: "OXE, DROID"   # or "none"
training_gpu_hours: 128

open_source: true
code_url: "https://github.com/you/your-method"
checkpoint_url: "https://huggingface.co/you/your-method"

Size limits

Keep the bundle under TBC GB. If your checkpoint is larger, host it separately and reference it from metadata.yaml rather than inlining it. Compress logs/ if needed — it is usually the largest part.

This file is produced by the harness. Do not edit it by hand — the verification step re-derives it from your logs and a mismatch invalidates the submission. It is documented here so you can confirm the harness ran over everything it should have.

{
  "roboeval_release": "corl2026-challenge",
  "method_name": "YourMethod-v2",
  "episodes_per_variation": 50,
  "seeds": [0, 1, 2, "..."],

  "results": {
    "LiftPot": {
      "Static": {
        "episodes": 50,
        "metrics": {
          "success": { "mean": 0.62, "std": 0.11 },
          "tp":      { "mean": 0.78, "std": 0.07 },
          "bavd":    { "mean": 0.19, "std": 0.03 },
          "bgvd":    { "mean": 0.08, "std": 0.01 },
          "cpl":     { "mean": 0.94, "std": 0.15 },
          "ct":      { "mean": 4.81, "std": 0.62 },
          "ecc":     { "mean": 1.90, "std": 0.44 },
          "jpl":     { "mean": 5.72, "std": 0.98 },
          "mcj":     { "mean": 5.11, "std": 0.87 },
          "mjj":     { "mean": 27.4, "std": 3.90 },
          "opl":     { "mean": 3.66, "std": 0.71 },
          "scc":     { "mean": 1.12, "std": 0.51 },
          "sc":      { "mean": 0.18, "std": 0.10 },
          "tl":      { "mean": 601.3, "std": 84.2 }
        }
      },
      "Position":                { "...": "..." },
      "Orientation":             { "...": "..." },
      "PositionAndOrientation":  { "...": "..." }
    },

    "LiftTray":                { "...": "..." },
    "StackTwoBlocks":          { "...": "..." },
    "StackSingleBookShelf":    { "...": "..." },
    "CubeHandover":            { "...": "..." },
    "PackBox":                 { "...": "..." },
    "PickSingleBookFromTable": { "...": "..." },
    "RotateValve":             { "...": "..." }
  }
}

Metric keys match the leaderboard schema in static/data/leaderboard_schema.json. Every task family and every variation must be present; a missing entry is scored as zero outcome for that variation rather than excluded from the average.

A short report is required for every submission — maximum 4 pages plus references, CoRL format. It is not peer-reviewed, and it does not affect your ranking. It exists so the community can learn something from the leaderboard beyond a row of numbers. Cover:

  • Method. Architecture, training objective, and what is novel relative to the baselines.
  • Data. Which demonstrations you trained on, any data you collected or generated, and any external pretraining.
  • Interface. Action mode, observation mode, control frequency, and any test-time compute.
  • Diagnostic analysis. Where your policy fails and why, using the stage-level and behavioral metrics. This is the part we most want to read.
  • Compute. Training GPU-hours and hardware.

Reports can become workshop papers

Challenge reports are welcome as submissions to the workshop's call for papers (short papers up to 5 pages, double-blind via OpenReview). See the workshop site for those deadlines, which are separate from the challenge deadline.

Submissions are made through a Google Form. One entry per submission; the most recent valid entry is the one ranked.

Submission form opens at challenge launch

The form link will be published here and on the workshop site. Until then, direct questions to yiruwang@cs.washington.edu.

The form collects:

  • Team name, contact email, and affiliation
  • Method name as it should appear on the leaderboard
  • Download link to your submission bundle
  • Whether you are opting into the open-source award
  • Confirmation that you have read and followed the rules

You receive an acknowledgment within TBC business days. Leaderboard entries appear after the automated schema check passes.

Every submission goes through two checks, and leading submissions go through a third.

1. Schema check Automated. Confirms results.json is well-formed, covers all 28 variations, and is consistent with the per-episode logs. Runs on submission; failures are reported back so you can fix and resubmit.
2. Consistency check Automated. Recomputes metrics from logs/ and compares against the reported aggregates. Divergence beyond floating-point tolerance is flagged for manual review.
3. Held-out re-run Manual, for prize-eligible entries. Organizers load your checkpoint and re-evaluate on an unpublished seed set. Results must hold up within normal run-to-run variance.

If step 3 cannot be completed — the checkpoint does not load, dependencies are unresolvable, or the inference code is incomplete — organizers contact you once with a fixed window to supply a working artifact. Entries that remain unreproducible are not eligible for prizes, though they may still be listed on the leaderboard marked as unverified.

Do I have to train on the provided demonstrations?

No. The 3,000+ demonstrations are provided as a convenience. You may train from scratch, use your own data, or fine-tune a pretrained model. Just declare what you used.

Can I submit a separate model per task?

Yes, as long as one submission bundle covers all 28 variations and the report says so. A generalist policy is not required, though we expect it to be discussed in the report.

Can I use ground-truth object poses?

Only if you declare the lightweight observation mode, which exposes joint positions and object poses through the standard interface. Reading state the mode does not expose is a rules violation. Submissions are grouped by observation mode on the leaderboard so image-based and state-based methods are comparable.

My policy runs slower than 20 Hz. Is that a problem?

No. Declare your actual control frequency in metadata.yaml. The simulator steps at the declared rate, so a slower policy does not get free extra time within an episode, but it is not disqualified either. Completion Time and Trajectory Length will reflect the difference.

Can I update my submission after the deadline?

Only to fix a reproducibility problem the organizers raise during verification. You cannot change the policy or improve results after the deadline.

Is a workshop paper required to compete?

No. The 4-page technical report is required; a workshop paper is optional and follows the workshop's own OpenReview deadlines.

Do I need to attend CoRL 2026 to win?

No, but winning teams are invited to present at the workshop on November 12, 2026, and we strongly encourage attendance. Travel support may be available — see Prizes.

Where do I ask technical questions?

Open an issue on the RoboEval repository. Questions about rules or eligibility go to yiruwang@cs.washington.edu.