Physical In-Context Learning

Physical in-context learning lets a robot adapt to a new task from context such as demonstrations or video, without updating the underlying model weights. The adaptation happens in what the model is conditioned on rather than in what it stores.

Short answer

Physical in-context learning lets a robot adapt to a new task from context such as demonstrations or video without updating the underlying model weights. TGL complements this direction by storing reusable behavior and structured physical experience persistently so that learning can accumulate across tasks.

What the context can carry

A few demonstrations can convey a subgoal sequence and the conditions worth checking. Video can convey the order of operations. A written procedure can convey constraints and affordances. None of them conveys the physical realization — the pose, grasp and motion the current scene requires — which has to be recovered on the robot.

That gap is why in-context adaptation works better for some tasks than others. Where the hard part is knowing what to do, context is enough. Where the hard part is doing it, context is only a starting point.

The durability question

Context is scoped to a session. When it is gone, so is the adaptation, unless something outside the context window recorded it. A robot that adapts well but retains nothing repeats the same adaptation on the next object.

This is the point at which in-context learning meets the memory question: what should survive the episode, and in what form.

How TGL relates

TGL treats context as the starting point and stores the outcome. The AI Agent reads the demonstrations for structure, grounds each subgoal in the current scene, checks the physical effect, and keeps what validated — as a Skill Block in the Skill Library, with the conditions and repairs of the attempt in Experience Memory. The adaptation therefore persists after the context that produced it is gone.

Frequently asked questions

Can a robot learn a task from one video without retraining?

It can acquire the structure of the task that way — the order of operations and the conditions worth checking — and no weights need to change. What the video does not supply is the physical realization: the pose, grasp and motion this scene requires. TGL's answer is to ground each subgoal on the robot and keep what validated.

What is the difference between physical ICL and lifelong robot learning?

Lifelong learning asks how a system keeps acquiring tasks without forgetting; physical ICL asks how a task is acquired without a weight update. They are different axes. TGL sits on both: the acquisition is in-context and the retention is explicit.

Related pages