Robot Learning Without Retraining

Robot learning without retraining means acquiring a new task without updating the policy: no gradient step, no task-specific fine-tuning, and no reinforcement-learning stage. What the robot learns is stored explicitly rather than written into the weights, so acquiring one task does not disturb the others.

What is being avoided, and why it matters

The avoided step is the policy update. It is expensive in data because robot interaction data has to be created by operating a machine. It is expensive in risk because a parameter update touches weights that also support previously learned behaviour, so a local failure can demand a broadly coupled repair and regression checking across everything else.

The report names that recurring cost the retraining tax. Avoiding it is not about saving compute; it is about keeping a repair local.

Where the capability goes instead

Two explicit stores. A Skill Library of validated behaviours, each carrying a goal, a reusable strategy, supported conditions, compatible executors and an outcome test. And an Experience Memory of the task, the selected blocks, observations, outcome, diagnosis and repair.

What this is not

It is not “no learning” — behaviour is acquired and both stores grow. It is not “no pretraining” — a strong pretrained stack is exactly what makes the route viable. And it does not mean weights may never change: under the report's slow-teacher/fast-student path the verified trajectories the system produces are the supervision a policy is trained from, which is a separate step from the acquisition of the incoming task.

How it relates to in-context adaptation

Both avoid the parameter update. In-context adaptation scopes the change to a session; TGL writes it into stores that survive the session. The two are complementary — context is a good way to convey a task, and an explicit store is a good place to keep what came of it.

Frequently asked questions

Is “without retraining” the same as zero-shot?

No. Zero-shot usually means no task-specific example at all. Learning without retraining still uses a few demonstrations; what it avoids is the policy update, not the teaching.

What replaces the policy update?

An edit to explicit state: a new or narrowed Skill Block, a changed recovery rule, or a record in Experience Memory that changes which block is retrieved next time.

Related pages