Free for humans

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

Pretty video generators do not finish a recipe: they must choose the next action from what they just made, do it, and stop. WorldGuide closes that loop with a planner and executor trained on the same step-level videos, beating a strong video model on two procedural benches even when the baseline gets reference plans.

arXiv:2610.124595 min readScore 73/100 · editorial triage · not peer reviewPaper hub2026-W42

The 30-second take

  • What: WorldGuide treats procedural video generation as closed-loop task execution in visual world space: from an initial image and a goal it predicts an atomic action, generates the clip, and uses that result to pick the next action or halt, with hierarchical visual memory and a new 59K-step, 245-task benchmark.
  • Why it matters: Long-horizon how-to intelligence — tutoring, assembly guidance, procedural planning — is still a scarce human service. A generator that can execute and stop is a step toward that assistance as default software, if success rates leave the 33–48% bench range and enter reliable use.
  • Who should care: Video-world-model and VLA researchers, benchmark builders, and teams who need goal-conditioned procedural generation rather than open-loop clips.

What the paper actually did

Video generators and video world models can synthesize plausible visual trajectories, but long-horizon procedural tasks need generation that adapts to what was actually produced: decide the next action from the generated state, execute it, and notice completion. Open-loop generation cannot adapt; existing closed-loop systems often lean on pretrained executors or indirect verification, leaving a gap between deciding an action and realizing it. The authors formulate procedural video generation as closed-loop task execution in visual world space and introduce WorldGuide. Given only an initial image and a task goal, it predicts an atomic action, generates the corresponding video clip, and uses that result to choose the next action or terminate. Planner and Executor train on the same step-level procedural demonstrations: the Planner predicts the next atomic action or completion from visual progress; the Executor is trained to realize predicted actions. Hierarchical visual memory keeps state across long horizons with bounded history-token cost. Because joint planner-executor training lacked step-level action-video supervision, they introduce WorldGuide Bench: about 59K step-annotated videos, 245 tasks, 27 procedural categories. Reported Task Success is 33.33% on WorldGuide-Bench versus 29.90% for MiniMax-H3 even when that model receives reference action plans, and 47.69% versus 32.73% on VideoCraft-Bench under goal-only conditioning.

What makes this disruptive

The scarce capability is finishing a procedure in generated video, not making a pretty clip. Coupling a learned planner to a learned executor, with memory that does not grow without bound, attacks the decide-vs-do gap the abstract calls out. Beating a strong recent video model that is handed reference plans is the rhetorical spike. Scarcity under pressure: expert procedural guidance and analysis. 33% task success is also an honest tell: this is a research loop, not a reliable how-to agent. The bench itself (59K steps) may matter as much as the model.

Why it matters (outside the lab)

Abundance lens: step-by-step visual help is still a specialist or expensive-staff service. If video models can plan, execute, and stop from a goal and a start image, that assistance starts looking like a software default. Near-term, use the bench and the gap to MiniMax-H3 to update roadmaps. Medium-term, cost and reliability decide whether anyone trusts generated procedures. No year. Do not confuse a 33–48% success paper with a universal tutor.

Limitations & open questions

Task Success remains well below reliable deployment. Gains over MiniMax-H3 are real in the abstract but modest on WorldGuide-Bench (33.33 vs 29.90). Generated video is not the same as acting on a robot or in the physical world. Hierarchical memory bounds tokens; it may still forget. Preprint ≠ product; abundance is not automatic. Read the PDF for action vocabularies, termination errors, and how much the new bench favors WorldGuide’s training format.

Explain ladder

Default article depth

Frame: closed-loop execution vs open-loop video. Two benches, one with the baseline given plans — that is the fair-play question. Ask what “atomic action” is and whether success is visual match or goal completion. Horizon: near for cognitive tools if reliability climbs; still research.

Key terms

Closed-loop generation
Each next action depends on the video (or state) just produced, rather than unrolling a plan blindly.
World model
A model that predicts how the visual world evolves given actions; here used to execute procedural goals.
WorldGuide Bench
The authors’ dataset of about 59K step-annotated videos across 245 tasks and 27 procedural categories.
Democratization of abundance
Editorial lens: turning scarce procedural expertise into cheaper default visual assistance if reliability rises.

Sources

Related explainers

Same topic and week first — keep exploring the scarcity → abundance map.

Editorial explainer · not peer review · always read the primary paper.

Byline: Disruptive Concepts editorial.