Methodology — the Xodexa Humanoid Index (XHI)
v2.0.0The XHI is a transparent, reproducible composite score from 0–100 for general-purpose humanoid robots. The design goal is to be defensible and hard to game: every input is an observable, published spec; every transform is documented here; and the pillar weights reflect where the field agrees value is actually created — autonomy and manipulation dominate, while raw locomotion, long since solved enough to be table stakes, no longer wins on its own.
Pillar weights
| Pillar | Weight | |
|---|---|---|
| Autonomy & Intelligence | 25% | |
| Manipulation & Dexterity | 20% | |
| Mobility & Locomotion | 15% | |
| Hardware & Engineering | 15% | |
| Commercial Readiness & Deployment | 15% | |
| Ecosystem & Viability | 10% |
Why these weights
- Autonomy & Intelligence — 25%. The single biggest open problem and value driver. A teleoperated robot is a puppet; a task-autonomous one is a worker.
- Manipulation & Dexterity — 20%. Useful work in human spaces is bottlenecked on hands.
- Mobility & Locomotion — 15%. Necessary but no longer the differentiator it was in the ASIMO era.
- Hardware & Engineering — 15%. DoF, actuation modernity, payload-to-weight, onboard compute.
- Commercial Readiness — 15%. Real paid pilots beat announcement videos.
- Ecosystem & Viability — 10%. Capital decides who survives to iterate — weighted lowest because money ≠ capability.
The anti-hype rule: teleoperation is capped
The most common way humanoid demos mislead is by showing a human-piloted robot as if it were autonomous. XHI scores the demonstrated autonomy mode on an explicit ladder, and teleoperation sits near the floor. This is why a polished consumer robot that relies on remote VR operators ranks below a plainer machine that genuinely does its own work.
| Autonomy level (base score) | Points |
|---|---|
| research | 20 |
| teleoperated | 35 |
| supervised-autonomy | 70 |
| task-autonomous | 92 |
Bonuses (deliberately narrow, so generic marketing language can't trigger them): +8 for a named VLA/foundation-policy system (e.g. GR00T, Helix, RT-2) or a specific published technique (e.g. diffusion policy, world model); +5 for a named onboard compute platform (e.g. Jetson Thor, Qualcomm Dragonwing). Capped at 100.
How each pillar is computed
- Autonomy = autonomy-ladder base + AI/compute bonuses.
- Manipulation = hand-DoF (55%, vs a 22-DoF human-hand reference) + payload (30%) + dexterous-hands flag (15%).
- Mobility = walk speed (55%, vs 2.5 m/s) + runtime (45%, vs 5 h).
- Hardware = total DoF (50%, vs 60) + payload-to-weight efficiency (25%) + actuation modernity (electric > hybrid > hydraulic) + onboard compute.
- Commercial = maturity ladder + 5 pts per named deployment (max +20).
- Ecosystem = funding (60%) + valuation (40%), with a floor for robots whose parent company is publicly traded (Tesla, Xiaomi, XPeng, Honda, Samsung, LG, Hyundai/Boston Dynamics) — capital raised isn't a meaningful concept for a product line inside a public company, so a null "funding" field there isn't scored as zero.
A spec that isn't published or verifiable is excluded from its pillar's average rather than counted as zero — a robot with 2 of 3 known manipulation specs is scored on those 2, not diluted by the unknown third. This keeps thin data from being penalised the same as a demonstrated weakness.
Evidence coverage caps the claim. Exclusion alone would mean a pillar could report a perfect score off a single disclosed field, so withholding specs would pay better than publishing them. A pillar is therefore capped at 50 + 50 × the share of its evidence that exists: full disclosure can reach 100, a pillar resting on 15% of its evidence cannot exceed 57.5. It binds only where a high score rests on thin data. Confirming a spec always scores at or above staying silent about it, including when the confirmed value is unflattering.
Commercial maturity ladder
| Status (base score) | Points |
|---|---|
| research | 10 |
| retired | 15 |
| prototype | 28 |
| pilot | 52 |
| limited-production | 76 |
| commercial | 94 |
Corporate-backed viability floor applies to: Boston Dynamics, Honda, Hyundai, LG, Samsung, Tesla, XPeng, Xiaomi.
Normalisation reference points
Sub-metrics are min-max normalised against frontier reference values, so a perfect 100 means "at or beyond the best demonstrated humanoid", not merely best in this list. A spec that isn't published is excluded from its pillar's weighted average rather than scored as 0, and the pillar is then capped by how much of its evidence exists — see "How each pillar is computed" above.
app/ranking.py and via the methodology API, so anyone can re-weight and recompute.Frequently asked questions
How is the Xodexa Humanoid Index (XHI) calculated?
XHI is a weighted sum of six 0–100 pillar scores — autonomy & intelligence, manipulation & dexterity, mobility & locomotion, hardware & engineering, commercial readiness, and ecosystem & viability — each normalised from published specs, then combined as XHI = Σ weight × pillar_score.
Why is a teleoperated robot scored lower than an autonomous one?
XHI scores the demonstrated autonomy mode on an explicit ladder, and teleoperation sits near the floor of that ladder. This is the anti-hype rule: a polished consumer robot that relies on a remote human operator ranks below a plainer machine that genuinely does its own work.
Which pillar counts most toward a robot's XHI score?
Autonomy & Intelligence, at 25% — the single biggest open problem and value driver. Manipulation & Dexterity is next at 20%. Ecosystem & Viability is weighted lowest, at 10%, because capital decides who survives to iterate but money is not the same as capability.
How reliable is the underlying data?
Every spec is cross-checked against at least two independent reputable sources, and each robot carries a per-figure confidence grade (high, medium, or low). Where a vendor claim could not be independently corroborated, the conservative reading was used rather than the vendor's own figure.