Adapting a VLA Model Without Retraining

Vision-language-action (VLA) models normally absorb an unfamiliar task by collecting more robot data and updating the policy. Teach-and-Grow Learning asks a narrower question: what can still be adapted while the pretrained VLA weights stay frozen? Its answer is to store the new capability as an explicit Skill Block and leave the model alone.

What the pretrained model is still for

Freezing the weights does not make the model passive. The VLA stack still supplies the perception, language grounding and control priors that let the robot interpret a scene and move through it. What it does not supply is a place to put a new task, because in the end-to-end route the only such place is the parameters.

TGL adds a second place. The AI Agent composes subgoals into Skill Blocks, each grounded in the current observation and checked against an outcome test, and stores the ones that pass.

Where the adaptation actually happens

Three things change during acquisition, and none of them is a weight: the set of available Skill Blocks, the retrieval that selects among them, and the Experience Memory that records what happened. Adaptation is therefore a change in the explicit state the AI Agent reasons over, not a change in the model.

In a new scene the same block can produce a different physical realization, because object bindings, grasp geometry and collision-free motion are recomputed from what the robot currently observes.

What this does and does not buy

It buys locality: repairing one behaviour is an edit to one explicit object rather than a parameter update with regression risk across everything else. It does not buy unlimited capability — the frozen prior still bounds what the robot can perceive and do, and grounding can still fail. It also does not remove the need for validated scope: an overgeneralised block is a real failure mode, which is why candidates are tested beyond their teaching demonstrations before admission.

Relation to other adaptation routes

Prompting, in-context adaptation and parameter-efficient fine-tuning also try to avoid full retraining. They differ in where the adapted knowledge lives: in a context window, in a small set of adapter weights, or — in TGL's case — in an inspectable store of behaviours that a person can read, narrow or revert.

Related pages