Benchmarks standardized performance tests
3 protocols · 0 recorded resultsEvery score on this site is computed from published specifications — what a manufacturer discloses about its own robot, cross-checked against at least two independent sources. That makes the 175-robot index a disclosure measure, and the evidence-coverage cap is an explicit admission of the limit: a pillar resting on thin evidence may not claim the frontier.
A standardized apparatus, run by a neutral party, measures the thing those specifications only proxy for. This page registers the protocols that do that, and records results published under them. It is separate from the ranking on purpose — see Why nothing here is scored below.
Registered protocols
Humanoid Robot Baseline Performance Benchmark
The first standardized, apparatus-based performance benchmark proposed for humanoid robots since the 2015 DARPA Robotics Challenge. NIST is designing a low-footprint circuit of locomotion and manipulation tasks, drawn mostly from previously standardized NIST test methods with quantifiable performance metrics, to establish minimum performance expectations for commercially available humanoids across industrial, healthcare and home applications.
DARPA Robotics Challenge
The disaster-response robotics competition whose 2015 Finals were, until the NIST baseline benchmark, the last standardized performance evaluation of humanoid robots. NIST designed tests for the DRC in 2013–2014, and the NIST baseline benchmark is explicitly built on that work.
IEEE/RAS Legged Robot Challenge
An IEEE/RAS competition for legged robots whose existing rules and procedures the Humanoids 2026 Loco-Manipulation Challenge adapts rather than replaces.
Why nothing here is scored
No robot in this index has a published result under any registered protocol. The results ledger holds 0 records against 175 spec-scored robots, and that gap is the finding — it is the most accurate single statement available about the state of humanoid benchmarking today.
Adding a benchmark pillar to the index now would mean scoring 175 robots on evidence that exists for none of them — precisely the failure the coverage cap was built to stop. Whether measured results should move the composite is a methodology decision to take when there is real data to take it against, with a METHODOLOGY_VERSION bump and a regenerated golden fixture. Until then this registry records; it does not rank.
Evidence tiers
The same distinction the rest of the site draws between a vendor claim and an independently corroborated fact, applied to measurement rather than specification.
| Tier | Meaning |
|---|---|
measured-independent | Run on the standardized apparatus by the standards body or an accredited third-party facility. |
measured-vendor | Run on the standardized apparatus by the manufacturer and self-reported. |
claimed | The capability is claimed, but not measured on the standardized apparatus. |