Before you start
Read the rules first. The most common cause of an invalidated
submission is an edited environment or metrics file, which is easy to avoid and painful to discover after
the deadline.
01
Submission Workflow
-
Pin the challenge release
Check out the tagged challenge release so your numbers are comparable with everyone else's.
# TODO(organizers): replace with the real challenge tag at launch
git fetch --tags
git checkout corl2026-challenge
pip install -e ".[examples]"
-
Run the evaluation harness
Evaluate your policy across all 8 task families and 28 variations using the published evaluation seeds.
This writes the per-task, per-variation metrics that your ranking is computed from.
python examples/5_gather_metrics.py \
--policy path/to/your_policy.py \
--checkpoint path/to/checkpoint.pt \
--output submission/results.json
Exact flags follow the challenge release; --help is authoritative.
-
Assemble the submission bundle
Collect the checkpoint, metrics, rollout logs, and report into the layout described in
Submission Format.
-
Upload the bundle
Host it somewhere the organizers can download without an account request — Google Drive,
Hugging Face Hub, S3, or an institutional server. Hugging Face is preferred for checkpoints.
-
Submit the link through the form
One form entry per submission. See Submitting.
02
Your uploaded bundle must have this structure:
submission/
├── results.json # REQUIRED — harness output, see schema below
├── report.pdf # REQUIRED — technical report, max 4 pages + refs
├── metadata.yaml # REQUIRED — team and method declaration
├── checkpoint/ # REQUIRED — policy weights, loadable as declared
│ └── ...
├── policy/ # REQUIRED — inference code needed to load and run
│ ├── policy.py # the checkpoint against RoboEval
│ └── requirements.txt
└── logs/ # REQUIRED — raw per-episode evaluation logs
└── <Task>_<Variation>/
└── episode_<i>.json
metadata.yaml
team_name: "Your Team"
method_name: "YourMethod-v2"
contact_email: "you@example.edu"
affiliation: "Your Institution"
members:
- "First Author"
- "Second Author"
roboeval_release: "corl2026-challenge"
action_mode: "JointPosition(absolute=true, ee=false)"
observation_mode: "full" # full | lightweight
control_frequency: 20 # Hz
policy_seed: 0
training_data:
roboeval_demos: true # did you train on the provided demonstrations?
additional_demos: 0 # count of self-collected demonstrations
external_pretraining: "OXE, DROID" # or "none"
training_gpu_hours: 128
open_source: true
code_url: "https://github.com/you/your-method"
checkpoint_url: "https://huggingface.co/you/your-method"
Size limits
Keep the bundle under TBC GB. If your checkpoint is larger, host it
separately and reference it from metadata.yaml rather than inlining
it. Compress logs/ if needed — it is usually the largest part.
03
results.json Schema
This file is produced by the harness. Do not edit it by hand — the verification step re-derives it
from your logs and a mismatch invalidates the submission. It is documented here so you can confirm the
harness ran over everything it should have.
{
"roboeval_release": "corl2026-challenge",
"method_name": "YourMethod-v2",
"episodes_per_variation": 50,
"seeds": [0, 1, 2, "..."],
"results": {
"LiftPot": {
"Static": {
"episodes": 50,
"metrics": {
"success": { "mean": 0.62, "std": 0.11 },
"tp": { "mean": 0.78, "std": 0.07 },
"bavd": { "mean": 0.19, "std": 0.03 },
"bgvd": { "mean": 0.08, "std": 0.01 },
"cpl": { "mean": 0.94, "std": 0.15 },
"ct": { "mean": 4.81, "std": 0.62 },
"ecc": { "mean": 1.90, "std": 0.44 },
"jpl": { "mean": 5.72, "std": 0.98 },
"mcj": { "mean": 5.11, "std": 0.87 },
"mjj": { "mean": 27.4, "std": 3.90 },
"opl": { "mean": 3.66, "std": 0.71 },
"scc": { "mean": 1.12, "std": 0.51 },
"sc": { "mean": 0.18, "std": 0.10 },
"tl": { "mean": 601.3, "std": 84.2 }
}
},
"Position": { "...": "..." },
"Orientation": { "...": "..." },
"PositionAndOrientation": { "...": "..." }
},
"LiftTray": { "...": "..." },
"StackTwoBlocks": { "...": "..." },
"StackSingleBookShelf": { "...": "..." },
"CubeHandover": { "...": "..." },
"PackBox": { "...": "..." },
"PickSingleBookFromTable": { "...": "..." },
"RotateValve": { "...": "..." }
}
}
Metric keys match the leaderboard schema in
static/data/leaderboard_schema.json. Every task family and every
variation must be present; a missing entry is scored as zero outcome for that variation rather than
excluded from the average.
04
Technical Report
A short report is required for every submission — maximum 4 pages plus references, CoRL format.
It is not peer-reviewed, and it does not affect your ranking. It exists so the community can learn
something from the leaderboard beyond a row of numbers. Cover:
- Method. Architecture, training objective, and what is novel relative to the
baselines.
- Data. Which demonstrations you trained on, any data you collected or generated, and
any external pretraining.
- Interface. Action mode, observation mode, control frequency, and any test-time
compute.
- Diagnostic analysis. Where your policy fails and why, using the stage-level and
behavioral metrics. This is the part we most want to read.
- Compute. Training GPU-hours and hardware.
Reports can become workshop papers
Challenge reports are welcome as submissions to the workshop's call for papers (short papers up to 5
pages, double-blind via OpenReview). See the
workshop
site for those deadlines, which are separate from the challenge deadline.
05
Submitting
Submissions are made through a Google Form. One entry per submission; the most recent valid entry is the
one ranked.
The form collects:
- Team name, contact email, and affiliation
- Method name as it should appear on the leaderboard
- Download link to your submission bundle
- Whether you are opting into the open-source award
- Confirmation that you have read and followed the rules
You receive an acknowledgment within TBC business days. Leaderboard entries
appear after the automated schema check passes.
06
Verification
Every submission goes through two checks, and leading submissions go through a third.
| 1. Schema check |
Automated. Confirms results.json is well-formed, covers all 28
variations, and is consistent with the per-episode logs. Runs on submission; failures are reported
back so you can fix and resubmit. |
| 2. Consistency check |
Automated. Recomputes metrics from logs/ and compares against
the reported aggregates. Divergence beyond floating-point tolerance is flagged for manual review. |
| 3. Held-out re-run |
Manual, for prize-eligible entries. Organizers load your checkpoint and re-evaluate on an
unpublished seed set. Results must hold up within normal run-to-run variance. |
If step 3 cannot be completed — the checkpoint does not load, dependencies are unresolvable, or the
inference code is incomplete — organizers contact you once with a fixed window to supply a working
artifact. Entries that remain unreproducible are not eligible for prizes, though they may still be listed
on the leaderboard marked as unverified.
07
FAQ
Do I have to train on the provided demonstrations?
No. The 3,000+ demonstrations are provided as a convenience. You may train from scratch, use your own
data, or fine-tune a pretrained model. Just declare what you used.
Can I submit a separate model per task?
Yes, as long as one submission bundle covers all 28 variations and the report says so. A generalist
policy is not required, though we expect it to be discussed in the report.
Can I use ground-truth object poses?
Only if you declare the lightweight observation mode, which exposes
joint positions and object poses through the standard interface. Reading state the mode does not expose
is a rules violation. Submissions are grouped by observation mode on the leaderboard so image-based and
state-based methods are comparable.
My policy runs slower than 20 Hz. Is that a problem?
No. Declare your actual control frequency in metadata.yaml. The
simulator steps at the declared rate, so a slower policy does not get free extra time within an episode,
but it is not disqualified either. Completion Time and Trajectory Length will reflect the difference.
Can I update my submission after the deadline?
Only to fix a reproducibility problem the organizers raise during verification. You cannot change the
policy or improve results after the deadline.
Is a workshop paper required to compete?
No. The 4-page technical report is required; a workshop paper is optional and follows the workshop's
own OpenReview deadlines.
Do I need to attend CoRL 2026 to win?
No, but winning teams are invited to present at the workshop on November 12, 2026, and we strongly
encourage attendance. Travel support may be available — see
Prizes.
Where do I ask technical questions?
Open an issue on the
RoboEval
repository. Questions about rules or eligibility go to
yiruwang@cs.washington.edu.