HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution
An 18-task humanoid tool-use benchmark plus 3.1k ToolBook demos shows a wide gap between picking the right tool and actually finishing the job — including on a real Unitree G1.
The 30-second take
- What: HumanoidToolBench jointly tests tool selection, manipulation, and needed locomotion on a humanoid, with 18 tasks, three scenarios, three execution levels, two tool-set modes, and a 3.1k-demonstration dataset.
- Why it matters: Hardware is advancing faster than evaluated tool-use competence; a shared benchmark makes the scarce skill — using tools to exceed a body’s built-in limits — cheaper to measure and improve.
- Who should care: Humanoid policy researchers, groups training on Unitree G1, and anyone comparing selection accuracy to full task success.
What the paper actually did
The authors argue that humanoids will need tools to exceed their inherent physical limits, which requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion. Existing benchmarks, they say, do not jointly evaluate those capabilities on a humanoid. HumanoidToolBench is an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, paired with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. They evaluate seven policies in simulation and three on the real robot, and report substantial gaps between selecting a suitable tool and completing the task. Focused probes of GR00T N1.7 show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are pointed to a project page.
What makes this disruptive
Tool use is how scarce physical capability gets leveraged — a humanoid without a good tool loop is stuck with its built-in body. A benchmark that jointly scores selection, manipulation, and mobile execution, with real G1 data, fills a measurement hole. The reported gap between “picked a tool” and “finished the task” is the useful negative result: selection accuracy is not success. The GR00T N1.7 probes (unseen tools; executing under unrelated instructions) name a concrete failure mode of current policies. Open code and 3.1k demos are the abundance move: shared evaluation plus data rather than another one-off lab task. This is a benchmark paper, not a solved tool-using humanoid.
Why it matters (outside the lab)
Abundance lens: reliable physical work that needs tools is still scarce human labor or capital equipment. If the field can see, on one scorecard, that models pick tools but fail the job, training effort can target the real bottleneck. Horizon is mid: reliability and safety decide whether tool-using humanoids become default capacity. Near-term, use the bench to stop over-claiming from selection metrics. Medium-term, broader tools and scenes decide if this becomes ordinary infrastructure.
Limitations & open questions
Eighteen tasks, three scenarios, and two tool-set modes are a designed slice, not the open world. Real-robot evaluation covers three policies; seven are simulation-only. ToolBook’s 3.1k demos mix sim and G1 — domain gap remains. The GR00T probes are focused, not a full survey of every humanoid VLA. “Substantial gaps” should be read from tables in the PDF, not as a single universal percentage. A benchmark can narrow research to its own 18 tasks. No timeline to commercial tool-using humanoids. Preprint.
Explain ladder
Default article depth
Use this as a scorecard paper. The joint claim is selection plus manipulation plus locomotion on a humanoid, with real G1 demos. The result to remember is the selection-versus-completion gap and the GR00T probes on unseen tools and unrelated instructions. Compare coverage to prior robot tool-use suites before you change a roadmap. Horizon: mid.
Key terms
- Tool-set mode
- A benchmark setting that changes which tools are available, used here to stress selection as well as execution.
- ToolBook
- The accompanying 3.1k-demonstration dataset collected in simulation and on a Unitree G1.
- Unitree G1
- A physical humanoid platform used for real-robot evaluation and demonstration collection in this paper.
- Democratization of abundance
- Editorial lens: measuring a scarce physical skill so it can become cheaper to train, without a fake ship date.
Sources
Related explainers
Same topic and week first — keep exploring the scarcity → abundance map.
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
2026-W41 · score 89 · Roboticssame weeksame topic
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
2026-W41 · score 87 · Roboticssame weeksame topic
SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation
2026-W41 · score 80 · Roboticssame weeksame topic
Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control
2026-W39 · score 93 · Roboticssame topic
Rapid Learning of Dexterous In-Hand Pen Writing through Real-Time Jacobian Estimation
2026-W38 · score 93 · Roboticssame topic
Disruptiveness
Editorial triage 0–100 · not peer review
- Novelty91
- Impact77
- Field heat94
- Practicality100
- Controversy34
