Training-Free Robot Learning
Training-free robot learning means acquiring a new robot capability without gradient updates, fine-tuning, or reinforcement learning. The pretrained model weights stay fixed, and whatever is learned about the new task is stored explicitly rather than written into the parameters. Teach-and-Grow Learning (TGL) is an architecture built on that constraint.
The precise definition
The term is narrower than it may sound. “Training-free” here describes the acquisition path, not the models: the pretrained AI Agent, perception and control components were trained by someone, and TGL assumes they are strong. What the term rules out is a gradient update, a task-specific fine-tuning run, or a reinforcement-learning stage when the robot meets a new task.
Where the new knowledge goes instead
If a capability is not written into weights, it has to live somewhere a person can inspect. TGL uses two explicit stores. The Skill Library holds validated behaviours — the goal, the reusable strategy, the supported conditions, compatible executors and an outcome test. The Experience Memory holds the context of use: which task, which blocks, what was observed, what happened, what the diagnosis was, and what repair was applied.
Why the constraint is interesting
Physical interaction data is expensive in a way text and code are not: it has to be created by operating a robot or a simulator. Because an end-to-end policy absorbs new behaviour into shared parameters, a local failure can demand a broadly coupled repair. Removing the parameter update from the acquisition path makes the update local — provided grounding, validation, compatibility and retrieval stay manageable, which is a condition the paper analyses rather than assumes.
What it is not
It is not “no learning”: behaviour is acquired, and the library and memory grow. It is not “no pretraining” — a strong prior is exactly what makes the route viable. And it is not a claim that parameters should never be touched: under the slow-teacher/fast-student split the verified trajectories this architecture produces are the supervision a policy is trained from, which is a separate step from the training-free acquisition of the incoming task.