SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation
SkeleWAM is a 57.1M-parameter world-action model that plans from a sparse 3D skeleton of joints, object centers, and contact points — 85.9% on LIBERO-Plus, 3.7 points above Cosmos-Policy.
The 30-second take
- What: The authors replace video or visual-latent world models with an online sparse 3D skeleton and train a compact model to generate actions and future skeletons from that geometry plus language.
- Why it matters: Pixel world models are heavy and store appearance that control does not need; a tiny geometric state is a path toward cheaper, more default robot policies.
- Who should care: Manipulation-policy groups, on-robot compute budgeters, and anyone comparing world-action models on LIBERO-Plus.
What the paper actually did
World action models jointly generate robot actions and predict future state. Existing WAMs, the authors say, typically predict videos or learned visual latents, which hide interaction geometry and may keep appearance unrelated to control. SkeleWAM represents a manipulation scene as a sparse 3D skeleton of robot joints, object centers, and interaction points, built online from current RGB-D observations and proprioception. That skeleton is a unified geometric state for action generation and future-skeleton prediction. Future-skeleton prediction supplies geometric supervision without visual reconstruction. At inference, actions come directly from the current skeleton and a language instruction; Medoid Action Consensus (MAC) is an auxiliary consensus over stochastic action samples. On LIBERO-Plus, SkeleWAM reaches 85.9% overall success with 57.1 million parameters, 3.7 percentage points above Cosmos-Policy.
What makes this disruptive
The scarce resource is on-robot compute and clean state for contact-rich control. Dropping pixels for a few 3D points is a bet that geometry is the world model control actually needs. A 57.1M model beating a named baseline by 3.7 points on LIBERO-Plus is a compactness claim, not a giant-scale claim. Online construction from RGB-D plus proprioception keeps the method deployable rather than oracle-mesh-only. MAC is a practical inference extra for stochastic samples. This is one benchmark and one skeleton design — still a real shift in what a WAM is allowed to predict.
Why it matters (outside the lab)
Abundance lens: reliable manipulation still requires scarce compute and engineering. If a sparse skeleton WAM can match or beat heavier video world models, capable policies get cheaper to run. Horizon is mid: reliability and unit economics decide defaults. Near-term, treat LIBERO-Plus 85.9% / 57.1M as a baseline to beat. Medium-term, more sensors, objects, and lighting decide whether skeletons become ordinary state.
Limitations & open questions
LIBERO-Plus is the reported arena; the abstract does not claim household-scale generality. A 3.7-point edge over Cosmos-Policy should be checked for training data, eval protocol, and parameter match. Skeleton extraction from RGB-D can fail when depth or object centers are wrong — the abstract does not quantify that perception error. Future-skeleton supervision is not the same as proving the predicted skeleton is a calibrated world model. MAC is auxiliary; the main inference path is current skeleton plus language. No product date. Preprint.
Explain ladder
Default article depth
The idea is “world action model, but the world is a stick figure.” Hold onto 85.9% LIBERO-Plus, 57.1M parameters, and +3.7 versus Cosmos-Policy. Ask whether your robots have the RGB-D and proprioception the online skeleton needs. Horizon: mid.
Key terms
- World action model (WAM)
- A policy that both outputs robot actions and predicts future state, here a future skeleton rather than video.
- Sparse 3D skeleton
- A compact geometric scene made of joints, object centers, and interaction points instead of dense pixels.
- Medoid Action Consensus (MAC)
- An auxiliary way to combine several stochastic action samples into one consensus action.
- LIBERO-Plus
- The manipulation benchmark on which SkeleWAM reports 85.9% overall success.
Sources
Related explainers
Same topic and week first — keep exploring the scarcity → abundance map.
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
2026-W41 · score 89 · Roboticssame weeksame topic
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
2026-W41 · score 87 · Roboticssame weeksame topic
HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution
2026-W41 · score 83 · Roboticssame weeksame topic
Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control
2026-W39 · score 93 · Roboticssame topic
Rapid Learning of Dexterous In-Hand Pen Writing through Real-Time Jacobian Estimation
2026-W38 · score 93 · Roboticssame topic
Disruptiveness
Editorial triage 0–100 · not peer review
- Novelty94
- Impact88
- Field heat77
- Practicality75
- Controversy45
