Free for humans

SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation

SkeleWAM is a 57.1M-parameter world-action model that plans from a sparse 3D skeleton of joints, object centers, and contact points — 85.9% on LIBERO-Plus, 3.7 points above Cosmos-Policy.

arXiv:2610.021205 min readScore 80/100 · editorial triage · not peer reviewPaper hub2026-W41

The 30-second take

  • What: The authors replace video or visual-latent world models with an online sparse 3D skeleton and train a compact model to generate actions and future skeletons from that geometry plus language.
  • Why it matters: Pixel world models are heavy and store appearance that control does not need; a tiny geometric state is a path toward cheaper, more default robot policies.
  • Who should care: Manipulation-policy groups, on-robot compute budgeters, and anyone comparing world-action models on LIBERO-Plus.

What the paper actually did

World action models jointly generate robot actions and predict future state. Existing WAMs, the authors say, typically predict videos or learned visual latents, which hide interaction geometry and may keep appearance unrelated to control. SkeleWAM represents a manipulation scene as a sparse 3D skeleton of robot joints, object centers, and interaction points, built online from current RGB-D observations and proprioception. That skeleton is a unified geometric state for action generation and future-skeleton prediction. Future-skeleton prediction supplies geometric supervision without visual reconstruction. At inference, actions come directly from the current skeleton and a language instruction; Medoid Action Consensus (MAC) is an auxiliary consensus over stochastic action samples. On LIBERO-Plus, SkeleWAM reaches 85.9% overall success with 57.1 million parameters, 3.7 percentage points above Cosmos-Policy.

What makes this disruptive

The scarce resource is on-robot compute and clean state for contact-rich control. Dropping pixels for a few 3D points is a bet that geometry is the world model control actually needs. A 57.1M model beating a named baseline by 3.7 points on LIBERO-Plus is a compactness claim, not a giant-scale claim. Online construction from RGB-D plus proprioception keeps the method deployable rather than oracle-mesh-only. MAC is a practical inference extra for stochastic samples. This is one benchmark and one skeleton design — still a real shift in what a WAM is allowed to predict.

Why it matters (outside the lab)

Abundance lens: reliable manipulation still requires scarce compute and engineering. If a sparse skeleton WAM can match or beat heavier video world models, capable policies get cheaper to run. Horizon is mid: reliability and unit economics decide defaults. Near-term, treat LIBERO-Plus 85.9% / 57.1M as a baseline to beat. Medium-term, more sensors, objects, and lighting decide whether skeletons become ordinary state.

Limitations & open questions

LIBERO-Plus is the reported arena; the abstract does not claim household-scale generality. A 3.7-point edge over Cosmos-Policy should be checked for training data, eval protocol, and parameter match. Skeleton extraction from RGB-D can fail when depth or object centers are wrong — the abstract does not quantify that perception error. Future-skeleton supervision is not the same as proving the predicted skeleton is a calibrated world model. MAC is auxiliary; the main inference path is current skeleton plus language. No product date. Preprint.

Explain ladder

Default article depth

The idea is “world action model, but the world is a stick figure.” Hold onto 85.9% LIBERO-Plus, 57.1M parameters, and +3.7 versus Cosmos-Policy. Ask whether your robots have the RGB-D and proprioception the online skeleton needs. Horizon: mid.

Key terms

World action model (WAM)
A policy that both outputs robot actions and predicts future state, here a future skeleton rather than video.
Sparse 3D skeleton
A compact geometric scene made of joints, object centers, and interaction points instead of dense pixels.
Medoid Action Consensus (MAC)
An auxiliary way to combine several stochastic action samples into one consensus action.
LIBERO-Plus
The manipulation benchmark on which SkeleWAM reports 85.9% overall success.

Sources

Related explainers

Same topic and week first — keep exploring the scarcity → abundance map.

Editorial explainer · not peer review · always read the primary paper.

Byline: Disruptive Concepts editorial.