ArenaGPT

BETA

Robot intelligence, compared.

Illustrative dataset. These six model profiles use the existing Arena seed scores. They are not independently measured results or deployment rankings.

About the data ↗

Physical Intelligence / GENERALIST VLA

π0.5Model analysis

Compare task-level scores, explore the cohort, and see what evidence is still missing.

MODEL COMPARISON

Overall index

Select a bar to inspect a model.

6 models

Illustrative index · 0–100 · Higher is better

π0.5 Comparison modelsShared zero baseline
TASK-LEVEL COMPARISONS

Different work. Different strengths.

The same selected cohort, separated by task.
All charts use a fixed 0–100 scale.

Manipulation

Illustrative scores · 6 models

Illustrative index · 0–100 · Higher is better

Navigation

Illustrative scores · 6 models

Illustrative index · 0–100 · Higher is better

Embodied reasoning

Illustrative scores · 6 models

Illustrative index · 0–100 · Higher is better

BEYOND TASK SCORES

What matters in the field.

Explore the HILO protocol ↗
Human BurdenNot measuredLower is better

Human attention, supervision, and recovery under matched task exposure.

MTHINot measuredLonger is better

Mean time to human intervention, reported with exposure and uncertainty.

Response latencyNot measuredLower is better

Measured decision and action latency, with p50, p95, and hardware disclosed.

Cost per missionNot measuredLower is better

Compute, energy, human time, maintenance, and recovery per useful outcome.

EVIDENCE & METHODOLOGY

Know what a bar can tell you.

ArenaGPT is Embodied Arena’s analysis interface. The current seed dataset demonstrates comparison behavior; it does not establish a winner or a new trained model.

Read WANTED-10K evidence requirements ↗
One shared dataset
These values also power the robotics leaderboard. Overall is a supplied seed index, not an average calculated from the three task scores.
Comparable evidence first
Measured rankings need a frozen protocol, matched tasks and hardware, sample counts, and uncertainty. Missing observations stay unmeasured.
Models, hardware, and safety
Jetson is an edge-compute platform. HILO evaluates the complete system; an independent safety kernel governs physical actions. Providers remain interchangeable.
ARENAGPT / EMBODIED ARENA

Analysis layout inspired by Artificial Analysis ↗. No Artificial Analysis benchmark scores are reproduced.