Humanoid Benchmark
Independent humanoid-robot index · Xodexa
XHI v2.0.0

Benchmarks standardized performance tests

3 protocols · 0 recorded results

Every score on this site is computed from published specifications — what a manufacturer discloses about its own robot, cross-checked against at least two independent sources. That makes the 175-robot index a disclosure measure, and the evidence-coverage cap is an explicit admission of the limit: a pillar resting on thin evidence may not claim the frontier.

A standardized apparatus, run by a neutral party, measures the thing those specifications only proxy for. This page registers the protocols that do that, and records results published under them. It is separate from the ranking on purpose — see Why nothing here is scored below.

Registered protocols

Humanoid Robot Baseline Performance Benchmark

NIST — Intelligent Systems Division · 2026
proposed 0 results

The first standardized, apparatus-based performance benchmark proposed for humanoid robots since the 2015 DARPA Robotics Challenge. NIST is designing a low-footprint circuit of locomotion and manipulation tasks, drawn mostly from previously standardized NIST test methods with quantifiable performance metrics, to establish minimum performance expectations for commercially available humanoids across industrial, healthcare and home applications.

1 · capability area
Domain-agnostic mobility and manipulation/dexterity
2 · capability area
Coordinated loco-manipulation
3 · capability area
Whole-body awareness
4 · capability area
Basic reasoning and scene understanding

DARPA Robotics Challenge

DARPA, with test methods designed by NIST · 2012–2015
historical 0 results

The disaster-response robotics competition whose 2015 Finals were, until the NIST baseline benchmark, the last standardized performance evaluation of humanoid robots. NIST designed tests for the DRC in 2013–2014, and the NIST baseline benchmark is explicitly built on that work.

IEEE/RAS Legged Robot Challenge

IEEE Robotics and Automation Society
active 0 results

An IEEE/RAS competition for legged robots whose existing rules and procedures the Humanoids 2026 Loco-Manipulation Challenge adapts rather than replaces.

Why nothing here is scored

No robot in this index has a published result under any registered protocol. The results ledger holds 0 records against 175 spec-scored robots, and that gap is the finding — it is the most accurate single statement available about the state of humanoid benchmarking today.

Adding a benchmark pillar to the index now would mean scoring 175 robots on evidence that exists for none of them — precisely the failure the coverage cap was built to stop. Whether measured results should move the composite is a methodology decision to take when there is real data to take it against, with a METHODOLOGY_VERSION bump and a regenerated golden fixture. Until then this registry records; it does not rank.

Evidence tiers

The same distinction the rest of the site draws between a vendor claim and an independently corroborated fact, applied to measurement rather than specification.

TierMeaning
measured-independentRun on the standardized apparatus by the standards body or an accredited third-party facility.
measured-vendorRun on the standardized apparatus by the manufacturer and self-reported.
claimedThe capability is claimed, but not measured on the standardized apparatus.
No metrics are invented here. NIST has published the four capability areas its baseline benchmark exercises, but not the finalized task list or the per-task metrics — those are stated for release in summer 2026. Where a figure has not been published, this registry records that it has not been published, rather than inferring one.