Free for humans

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

RPG improves a frozen robot execution stack by practicing in simulation and editing skills plus the system prompt — 28.6% to 95.0% on 22 held-out sim tasks, then 30/30 physical trials after calibration.

arXiv:2610.022045 min readScore 87/100 · editorial triage · not peer reviewPaper hub2026-W41

The 30-second take

  • What: Reconstruct, Practice, Go Real (RPG) finds manipulation skills in offline data, practices related tasks in simulation, and keeps only the skill and prompt edits that survive cross-task tests — without updating model weights.
  • Why it matters: Human effort to write skills, rewards, and perception–control glue is the scarce cost; if a frozen stack can self-improve from practice, more robot competence can become a default library.
  • Who should care: Robot-learning labs, companies sitting on demonstration logs, and teams who want improvement without another fine-tune cycle.

What the paper actually did

RPG is a framework for autonomous improvement of robot execution systems that does not update model weights. It identifies manipulation capabilities in an offline dataset and builds related practice tasks in simulation. During practice it uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. From those diagnoses it develops new reusable symbolic skills, refines existing skills, and revises the system prompt. Cross-task evaluation tests individual candidate changes and merged revisions before anything is retained. At test time a multimodal LLM uses the resulting prompt and skill library to coordinate perception and control. On held-out initializations of 22 manipulation tasks, success rises from 28.6% after the first practice round to 95.0% after 15 rounds, beating evaluated baselines including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a shared calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials — ten trials on each of three tasks.

What makes this disruptive

Weight updates are the default story for robot improvement. RPG instead treats skills and the system prompt as the writable surface, with simulation practice plus privileged diagnostics as the teacher. The reported jump from 28.6% to 95.0% across 15 rounds on 22 held-out initializations is a strong self-improvement curve against named baselines. The 30/30 real-robot result, after calibration and hardware adaptation, is the transfer claim. Because the underlying models stay frozen, the scarce resource becomes practice time and a retain/reject test, not another training run. That is a different abundance path: more competence from the same weights. It is still a research stack with a project page, not a general household robot.

Why it matters (outside the lab)

Abundance lens: reliable manipulation still needs scarce human skill-writing, reward design, and integration work. If frozen systems can reconstruct practice tasks and keep only edits that survive cross-task tests, that human loop gets cheaper. Horizon is mid: reliability, safety, and unit economics decide defaults. Near-term, the paper is a prompt-and-skill self-improvement recipe. Medium-term, whether 15 practice rounds and a calibration procedure generalize beyond 22 sim tasks and three physical tasks is the real gate.

Limitations & open questions

Physical success is 30 trials on three tasks after calibration and hardware adaptation — not an unbounded real-world suite. The 95.0% figure is on held-out initializations of 22 simulation tasks, not on an open-world home. Privileged simulator state is available during practice; that oracle is not present at real test time. Baseline names (ASPIRE, CaP-Agent0 / GPT-6 Astra Pro) should be checked in the PDF for fair compute and prompt budgets. Skill edits that pass cross-task eval can still overfit the practice distribution. No claim that skipping weight updates always beats fine-tuning. Preprint; independent replication is open.

Explain ladder

Default article depth

Separate three numbers: 28.6% → 95.0% over 15 sim practice rounds, baseline gaps (ASPIRE 75.5%, CaP-Agent0 60.0%), and 30/30 real trials on three tasks after calibration. The method claim is “no weight updates” — improvement lives in skills and the system prompt. Ask how much privileged sim state is doing during diagnosis. Horizon: mid, after broader hardware.

Key terms

Symbolic skill
A reusable, named manipulation routine stored in a library rather than baked only into neural weights.
System prompt
The standing instructions a multimodal model uses to coordinate perception and control at test time.
Privileged simulator state
Ground-truth simulation information available during practice but not to the deployed robot.
Democratization of abundance
Editorial lens: making scarce physical competence cheaper to grow, without inventing a product date.

Sources

Related explainers

Same topic and week first — keep exploring the scarcity → abundance map.

Editorial explainer · not peer review · always read the primary paper.

Byline: Disruptive Concepts editorial.