Free for humans

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

An 18-task humanoid tool-use benchmark plus 3.1k ToolBook demos shows a wide gap between picking the right tool and actually finishing the job — including on a real Unitree G1.

arXiv:2610.020895 min readScore 83/100 · editorial triage · not peer reviewPaper hub2026-W41

The 30-second take

  • What: HumanoidToolBench jointly tests tool selection, manipulation, and needed locomotion on a humanoid, with 18 tasks, three scenarios, three execution levels, two tool-set modes, and a 3.1k-demonstration dataset.
  • Why it matters: Hardware is advancing faster than evaluated tool-use competence; a shared benchmark makes the scarce skill — using tools to exceed a body’s built-in limits — cheaper to measure and improve.
  • Who should care: Humanoid policy researchers, groups training on Unitree G1, and anyone comparing selection accuracy to full task success.

What the paper actually did

The authors argue that humanoids will need tools to exceed their inherent physical limits, which requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion. Existing benchmarks, they say, do not jointly evaluate those capabilities on a humanoid. HumanoidToolBench is an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, paired with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. They evaluate seven policies in simulation and three on the real robot, and report substantial gaps between selecting a suitable tool and completing the task. Focused probes of GR00T N1.7 show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are pointed to a project page.

What makes this disruptive

Tool use is how scarce physical capability gets leveraged — a humanoid without a good tool loop is stuck with its built-in body. A benchmark that jointly scores selection, manipulation, and mobile execution, with real G1 data, fills a measurement hole. The reported gap between “picked a tool” and “finished the task” is the useful negative result: selection accuracy is not success. The GR00T N1.7 probes (unseen tools; executing under unrelated instructions) name a concrete failure mode of current policies. Open code and 3.1k demos are the abundance move: shared evaluation plus data rather than another one-off lab task. This is a benchmark paper, not a solved tool-using humanoid.

Why it matters (outside the lab)

Abundance lens: reliable physical work that needs tools is still scarce human labor or capital equipment. If the field can see, on one scorecard, that models pick tools but fail the job, training effort can target the real bottleneck. Horizon is mid: reliability and safety decide whether tool-using humanoids become default capacity. Near-term, use the bench to stop over-claiming from selection metrics. Medium-term, broader tools and scenes decide if this becomes ordinary infrastructure.

Limitations & open questions

Eighteen tasks, three scenarios, and two tool-set modes are a designed slice, not the open world. Real-robot evaluation covers three policies; seven are simulation-only. ToolBook’s 3.1k demos mix sim and G1 — domain gap remains. The GR00T probes are focused, not a full survey of every humanoid VLA. “Substantial gaps” should be read from tables in the PDF, not as a single universal percentage. A benchmark can narrow research to its own 18 tasks. No timeline to commercial tool-using humanoids. Preprint.

Explain ladder

Default article depth

Use this as a scorecard paper. The joint claim is selection plus manipulation plus locomotion on a humanoid, with real G1 demos. The result to remember is the selection-versus-completion gap and the GR00T probes on unseen tools and unrelated instructions. Compare coverage to prior robot tool-use suites before you change a roadmap. Horizon: mid.

Key terms

Tool-set mode
A benchmark setting that changes which tools are available, used here to stress selection as well as execution.
ToolBook
The accompanying 3.1k-demonstration dataset collected in simulation and on a Unitree G1.
Unitree G1
A physical humanoid platform used for real-robot evaluation and demonstration collection in this paper.
Democratization of abundance
Editorial lens: measuring a scarce physical skill so it can become cheaper to train, without a fake ship date.

Sources

Related explainers

Same topic and week first — keep exploring the scarcity → abundance map.

Editorial explainer · not peer review · always read the primary paper.

Byline: Disruptive Concepts editorial.