Humanoid Benchmark
Independent humanoid-robot index · Xodexa
XHI v2.0.0
Benchmarks / NIST Baseline

Humanoid Robot Baseline Performance Benchmark

NIST — Intelligent Systems Division
proposed 2026 USA confidence: high as of 2026-08

The first standardized, apparatus-based performance benchmark proposed for humanoid robots since the 2015 DARPA Robotics Challenge. NIST is designing a low-footprint circuit of locomotion and manipulation tasks, drawn mostly from previously standardized NIST test methods with quantifiable performance metrics, to establish minimum performance expectations for commercially available humanoids across industrial, healthcare and home applications.

Why it matters to this index

Every score on this site is computed from published specifications — what a manufacturer discloses about its own robot, cross-checked against independent sources. That makes the index a disclosure measure, and its evidence-coverage cap is an explicit admission of the limit. A standardized apparatus run by a neutral party measures the thing the specifications only proxy for. When results from this benchmark exist, they are a stronger class of evidence than anything the index currently scores, and this page is where they will be recorded.

Capability areas

#AreaWhat it exercises
1Domain-agnostic mobility and manipulation/dexterityBasic locomotion and hand/arm capability measured independently of any single application domain.
2Coordinated loco-manipulationMobility and manipulation exercised together rather than in isolation — carrying, moving while holding, manipulating from an unstable base.
3Whole-body awarenessConfined-space tasks that require the robot to reason about the position and clearance of its entire body, not just its end effectors.
4Basic reasoning and scene understandingMinimal perception and reasoning about the scene the tasks are set in.
Metrics not yet published. NIST has published the four capability areas but not the finalized task list or the per-task metrics. Task lists, 3D models, fabrication plans, proposal templates, procedures and rules are stated to be released in summer 2026. No metric values, thresholds or scoring weights are recorded here because none have been published — this record will be extended when they are, not filled in by inference.

Apparatus

NIST will fabricate a limited number of apparatuses and distribute them free of charge to participating U.S. humanoid robot manufacturers and established regional testing facilities. The designs and 3D models are to be published so the apparatus can be used as a physical and/or virtual testbed for robot training and control development.

Distribution
free-to-participants
Virtual testbed
published

How to take part

  • At NIST
  • At a participating regional testing facility
  • At the manufacturer's own site, on an apparatus received from NIST

Contacts: Dr. Benjamin Beiter (benjamin.beiter@nist.gov) · Dr. Kamel S. Saidi (kamel.saidi@nist.gov)

Results handling

Results are collected under data-sharing agreements that protect participants' intellectual property and attribution. NIST has not announced a public per-robot leaderboard.

Events

Humanoids 2026 Loco-Manipulation Challenge

2026-12-07 – 2026-12-09 · Santa Clara Convention Center, Santa Clara, CA, USA

The competition instantiation of the benchmark, held during the 25th IEEE-RAS International Conference on Humanoid Robots (6–9 December 2026). One practice day and two competition days for 10–15 teams, running several cross-cutting basic tasks arranged in a circuit. Selected top humanoids are to be tested before and after the competition so participants can benchmark against them. Rules and procedures are variations on the IEEE/RAS Legged Robot Challenge. Autonomy is incentivized, but teleoperation scores points separately so the event stays accessible.

DeadlineDate
Competition proposals due2026-08-03
Acceptance notification2026-09-21

Recorded results

0

No results have been published under this protocol for any robot in the index. This ledger records only results actually published under the protocol — never a vendor capability claim reinterpreted as a measurement, and never a figure inferred from a spec sheet. It stays empty until there is something real to put in it.

Builds on