Free for humans

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

InterEvolve lets a humanoid reuse a frozen controller on new loco-manipulation tasks by evolving staged reward programs at test time — in simulation and, after evolution, on a physical Unitree G1.

arXiv:2610.021965 min readScore 89/100 · editorial triage · not peer reviewPaper hub2026-W41

The 30-second take

  • What: The system writes and revises staged reward programs so a preexisting humanoid controller can solve tasks it was never trained on, without retraining the controller.
  • Why it matters: Retraining whole-body policies is scarce and expensive; if test-time program search can unlock unused motor competence, new physical skills get cheaper to specify.
  • Who should care: Humanoid researchers, sim-to-real control teams, and anyone watching whether language-model planners can retarget a frozen robot stack.

What the paper actually did

InterEvolve studies test-time evolution for humanoid loco-manipulation: solving tasks a controller was never trained for by repurposing existing skills, improving from its own attempts, and retaining what it learns, without retraining. The interface between planning and control is a pair of pieces. First, an object-aware forward-backward (FB) behavioral foundation model uses object residuals on a frozen body prior so a new reward about the body or objects becomes loco-manipulation behavior at test time. Second, tasks are specified as reward programs — staged rewards with completion conditions and tunable constants. A language-model agent revises program structure from execution feedback and a library of verified programs, while a numerical optimizer tunes constants. Candidates are verified across parallel simulation scenarios. The authors report that human-designed rewards leave much of the FB model’s competence unused, whereas evolved programs release it, sometimes via novel strategies. Evolved skills cover diverse tasks, complex scenes, and long-horizon compositions in simulation, and they run autonomously on a physical Unitree G1 from egocentric onboard perception.

What makes this disruptive

The usual scarcity in humanoid work is a new policy for every new contact-rich task. InterEvolve treats the controller as already containing unused competence and spends search on the reward program instead of weight updates. That is a different cost curve: LLM-guided structure search plus numerical constant tuning, verified in parallel simulation, then transfer of the evolved skill to a real G1. Claiming that human-written rewards leave competence “untapped” is a direct critique of the current specification interface. Physical, autonomous execution from egocentric onboard sensing is the part that keeps this from being only a simulator story. Still, the paper is about releasing a frozen model’s skills, not about proving a general humanoid product.

Why it matters (outside the lab)

Abundance lens: reliable physical work and whole-body loco-manipulation are still scarce human or capital-heavy capabilities. If a frozen foundation controller plus evolving reward programs can cover new tasks, the scarce thing becomes a good skill library rather than a full retrain. Horizon is mid: reliability, safety, and unit economics decide whether this is a lab demo or a default way to retarget humanoids. Near-term, use it to update how you specify tasks to behavioral foundation models. Medium-term, replication on more robots and scenes decides if test-time evolution becomes ordinary practice.

Limitations & open questions

The physical result is that evolved skills run on a Unitree G1 from egocentric perception; the abstract does not give a full real-world task table or success rates. Evolution and verification happen in simulation; sim-to-real gaps remain a gate. The method assumes a broad preexisting controller and an object-aware FB model — it does not create motor competence from scratch. Language-model program revision can invent strategies that look clever in sim and fail under real contact or perception noise. Parallel scenario verification reduces but does not eliminate overfitting to the simulator. No calendar claim that humanoid labor becomes cheap. Preprint, not a shipped skill store.

Explain ladder

Default article depth

Read this as a specification paper: the scarce interface is the reward program, not another end-to-end policy. Separate the simulation claim (evolved programs beat human-designed rewards on an FB controller) from the robot claim (autonomous G1 execution of evolved skills). Ask what “never trained for” means given the frozen body prior. Horizon: mid, after more hardware and safety evidence.

Key terms

Loco-manipulation
Robot tasks that mix walking or balancing with using the arms to contact and move objects.
Reward program
A staged, executable reward specification with completion checks and tunable constants, used here as the planner–controller interface.
Test-time evolution
Improving behavior on a new task by searching during deployment rather than retraining model weights.
Forward-backward (FB) model
A behavioral foundation model that, in this paper, turns a new body- or object-centric reward into loco-manipulation behavior.

Sources

Related explainers

Same topic and week first — keep exploring the scarcity → abundance map.

Editorial explainer · not peer review · always read the primary paper.

Byline: Disruptive Concepts editorial.