InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
InterEvolve lets a humanoid reuse a frozen controller on new loco-manipulation tasks by evolving staged reward programs at test time — in simulation and, after evolution, on a physical Unitree G1.
The 30-second take
- What: The system writes and revises staged reward programs so a preexisting humanoid controller can solve tasks it was never trained on, without retraining the controller.
- Why it matters: Retraining whole-body policies is scarce and expensive; if test-time program search can unlock unused motor competence, new physical skills get cheaper to specify.
- Who should care: Humanoid researchers, sim-to-real control teams, and anyone watching whether language-model planners can retarget a frozen robot stack.
What the paper actually did
InterEvolve studies test-time evolution for humanoid loco-manipulation: solving tasks a controller was never trained for by repurposing existing skills, improving from its own attempts, and retaining what it learns, without retraining. The interface between planning and control is a pair of pieces. First, an object-aware forward-backward (FB) behavioral foundation model uses object residuals on a frozen body prior so a new reward about the body or objects becomes loco-manipulation behavior at test time. Second, tasks are specified as reward programs — staged rewards with completion conditions and tunable constants. A language-model agent revises program structure from execution feedback and a library of verified programs, while a numerical optimizer tunes constants. Candidates are verified across parallel simulation scenarios. The authors report that human-designed rewards leave much of the FB model’s competence unused, whereas evolved programs release it, sometimes via novel strategies. Evolved skills cover diverse tasks, complex scenes, and long-horizon compositions in simulation, and they run autonomously on a physical Unitree G1 from egocentric onboard perception.
What makes this disruptive
The usual scarcity in humanoid work is a new policy for every new contact-rich task. InterEvolve treats the controller as already containing unused competence and spends search on the reward program instead of weight updates. That is a different cost curve: LLM-guided structure search plus numerical constant tuning, verified in parallel simulation, then transfer of the evolved skill to a real G1. Claiming that human-written rewards leave competence “untapped” is a direct critique of the current specification interface. Physical, autonomous execution from egocentric onboard sensing is the part that keeps this from being only a simulator story. Still, the paper is about releasing a frozen model’s skills, not about proving a general humanoid product.
Why it matters (outside the lab)
Abundance lens: reliable physical work and whole-body loco-manipulation are still scarce human or capital-heavy capabilities. If a frozen foundation controller plus evolving reward programs can cover new tasks, the scarce thing becomes a good skill library rather than a full retrain. Horizon is mid: reliability, safety, and unit economics decide whether this is a lab demo or a default way to retarget humanoids. Near-term, use it to update how you specify tasks to behavioral foundation models. Medium-term, replication on more robots and scenes decides if test-time evolution becomes ordinary practice.
Limitations & open questions
The physical result is that evolved skills run on a Unitree G1 from egocentric perception; the abstract does not give a full real-world task table or success rates. Evolution and verification happen in simulation; sim-to-real gaps remain a gate. The method assumes a broad preexisting controller and an object-aware FB model — it does not create motor competence from scratch. Language-model program revision can invent strategies that look clever in sim and fail under real contact or perception noise. Parallel scenario verification reduces but does not eliminate overfitting to the simulator. No calendar claim that humanoid labor becomes cheap. Preprint, not a shipped skill store.
Explain ladder
Default article depth
Read this as a specification paper: the scarce interface is the reward program, not another end-to-end policy. Separate the simulation claim (evolved programs beat human-designed rewards on an FB controller) from the robot claim (autonomous G1 execution of evolved skills). Ask what “never trained for” means given the frozen body prior. Horizon: mid, after more hardware and safety evidence.
Key terms
- Loco-manipulation
- Robot tasks that mix walking or balancing with using the arms to contact and move objects.
- Reward program
- A staged, executable reward specification with completion checks and tunable constants, used here as the planner–controller interface.
- Test-time evolution
- Improving behavior on a new task by searching during deployment rather than retraining model weights.
- Forward-backward (FB) model
- A behavioral foundation model that, in this paper, turns a new body- or object-centric reward into loco-manipulation behavior.
Sources
Related explainers
Same topic and week first — keep exploring the scarcity → abundance map.
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
2026-W41 · score 87 · Roboticssame weeksame topic
HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution
2026-W41 · score 83 · Roboticssame weeksame topic
SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation
2026-W41 · score 80 · Roboticssame weeksame topic
Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control
2026-W39 · score 93 · Roboticssame topic
Rapid Learning of Dexterous In-Hand Pen Writing through Real-Time Jacobian Estimation
2026-W38 · score 93 · Roboticssame topic
Disruptiveness
Editorial triage 0–100 · not peer review
- Novelty100
- Impact100
- Field heat100
- Practicality40
- Controversy46
