Overall index
Select a bar to inspect a model.
Illustrative index · 0–100 · Higher is better
Robot intelligence, compared.
Illustrative dataset. These six model profiles use the existing Arena seed scores. They are not independently measured results or deployment rankings.
About the data ↗Physical Intelligence / GENERALIST VLA
Compare task-level scores, explore the cohort, and see what evidence is still missing.
Select a bar to inspect a model.
Illustrative index · 0–100 · Higher is better
The same selected cohort, separated by task.
All charts use a fixed 0–100 scale.
Illustrative scores · 6 models
Illustrative index · 0–100 · Higher is better
Illustrative scores · 6 models
Illustrative index · 0–100 · Higher is better
Illustrative scores · 6 models
Illustrative index · 0–100 · Higher is better
Human attention, supervision, and recovery under matched task exposure.
Mean time to human intervention, reported with exposure and uncertainty.
Measured decision and action latency, with p50, p95, and hardware disclosed.
Compute, energy, human time, maintenance, and recovery per useful outcome.
ArenaGPT is Embodied Arena’s analysis interface. The current seed dataset demonstrates comparison behavior; it does not establish a winner or a new trained model.
Read WANTED-10K evidence requirements ↗Analysis layout inspired by Artificial Analysis ↗. No Artificial Analysis benchmark scores are reproduced.