Free for humans

ROWBench: Do Video Models Render What the Program Specifies?

PROWBench (presented as ROWBench) checks whether programmable world models actually depict the events a program logged — 170 constructed episodes, 600 proxy videos, including events off camera.

arXiv:2610.022055 min readScore 75/100 · editorial triage · not peer reviewPaper hub2026-W41

The 30-second take

  • What: The benchmark builds programmatic 3D episodes, logs entity states and timestamped events (even out of view), and scores whether generated video matches those executable records.
  • Why it matters: Pretty video is cheap; video that obeys an engine’s rules is still scarce — a shared test is how game-engine-like world models become measurable rather than vibes.
  • Who should care: World-model and game-engine researchers, video-generation teams, and anyone who needs logic-level fidelity rather than clip aesthetics.

What the paper actually did

The paper (title: ROWBench; abstract name: PROWBench) targets programmable world models that separate executable dynamics from visual generation. Existing benchmarks, the authors say, score visual quality, controllability, and instruction or physical adherence, but rarely fidelity to fine-grained, program-specified world events. PROWBench has 170 programmatically constructed episodes and 600 proxy videos across diverse scenes and interactions. It logs entity states and timestamped events, including those outside the camera’s field of view, as replayable world records, then renders synchronized views and proxy representations so generated videos can be checked against observable consequences of program execution. An extensible framework builds scenes, controls behaviors, and can render each camera in different representations such as coarse 3D and bounding boxes. Coverage includes first- and third-person views, with synchronized multi-view observations on a subset of episodes. Evaluation targets entity control, long-horizon memory, and two VLM-based metrics — Logic-Render Alignment and Interaction Success Rate — for timeline adherence and visual realization of timestamped engine events.

What makes this disruptive

Programmable world models are a bid to make game-engine logic a first-class object. Without a benchmark that keeps an off-camera ground-truth log, “the video followed the rules” stays untested. Logging events outside the field of view is the distinctive move: memory and hidden state become checkable. Two VLM metrics (Logic-Render Alignment, Interaction Success Rate) plus entity control and long-horizon memory give a scorecard beyond FID-like quality. Multi-view and proxy representations (coarse 3D, boxes) make the bench reusable. This is infrastructure for measurement, not a new generator that already passes.

Why it matters (outside the lab)

Abundance lens: accurate, rule-following interactive worlds are still expensive to build by hand. If generators must match a program’s timeline, cheaper default game-engine-like tools become a research target rather than a trailer. Horizon is mid: measurement first, products later. Near-term, use the bench to stop celebrating instruction adherence that ignores hidden events. Medium-term, broader scenes and non-VLM metrics decide if this is ordinary eval infrastructure.

Limitations & open questions

The public title says ROWBench while the abstract defines PROWBench — confirm the canonical name in the PDF. One hundred seventy episodes and 600 proxy videos are a designed corpus, not every game genre. Two headline metrics are VLM-based and inherit VLM error. Off-camera events are in the log; a generated video cannot depict them directly, so scoring “observable consequences” needs a careful definition. Multi-view exists only for a subset. No claim that current video models pass. Preprint.

Explain ladder

Default article depth

Read this as an evaluation design: executable world records versus rendered pixels. Note 170 episodes, 600 proxy videos, off-camera logged events, and the two VLM metrics. Do not treat the paper as proving today’s generators are faithful; it argues they have not been tested this way. Horizon: mid.

Key terms

Programmable world model
A system that keeps executable dynamics separate from the pixels it renders.
World record
A replayable log of entity states and timestamped events, including events off camera.
Logic-Render Alignment
A VLM-based score for whether the video follows the prescribed timeline of engine events.
Interaction Success Rate
A VLM-based score for whether timestamped interactions are visually realized.

Sources

Related explainers

Same topic and week first — keep exploring the scarcity → abundance map.

Editorial explainer · not peer review · always read the primary paper.

Byline: Disruptive Concepts editorial.