Teach and Grow is a training-free architecture for general robot learning. It sits inside embodied-intelligence and agentic-robotics research, next to vision-language-action (VLA) models and world-action models (WAM). The difference it proposes is not in how strong the underlying models are, but in where a newly acquired capability is stored.
Vision-language-action (VLA) models
VLA models map camera images, a language instruction and robot state directly to control, and they are the dominant architecture for general-purpose manipulation. They work because large multimodal pretraining produces representations that transfer. In their end-to-end route a new behaviour is absorbed into the model parameters through training: collect more robot data, optimise, re-check what the policy already supported. TGL keeps those models as the source of priors and moves new task knowledge somewhere else.
World-action models and learned dynamics
World-action models add learned physical dynamics, so the system reasons about how a scene will evolve rather than only reacting to its current state. This improves generalisation, at the cost of a heavier training cycle and a larger data requirement. The retraining burden TGL names applies to these models as much as to VLA policies: the repair path runs through the parameters.
Robot foundation models as the prior
TGL assumes a strong pretrained stack — a multimodal AI Agent for reasoning, plus specialist perception, grasping and motion tools. Robot foundation models supply exactly that prior. In the paper's framing they are held fixed while the explicit, inspectable stores grow: the Skill Library of validated behaviours and the Experience Memory of conditions and repairs.
Agentic robotics
The architecture is AI Agent-centered: the reasoning AI Agent reads the scene, chooses a subgoal and a tool, observes the result, and revises the remaining plan. This is the pattern that made tool-using language AI Agents useful, applied where a mistaken action changes the physical world. The AI Agent carries task-level reasoning; the robot-side executors carry geometry and continuous control.
Skill composition and lifelong learning
TGL's neighbours also include work on skill composition, skill libraries and lifelong or continual learning. The shared question is what a robot retains across tasks and how it is retrieved later. TGL's contribution is to make the retained objects explicit — a validated behaviour with a stated scope and outcome test — so that a person can narrow an overgeneralised skill, revise a recovery rule, or mark an executor version as incompatible.
Few-shot and sparse teaching
Teaching supplies the structure the AI Agent starts from: the subgoal sequence, the ordering, and the conditions worth checking. It does not supply the physical realization, which is recomputed for the current scene. That is why TGL can preserve the intended effect while the actual grasp, path and contact point differ from the teacher's.
What TGL keeps from these directions
A learned policy can still implement a Skill Block, execute a familiar composition, or serve as the fast student that takes over mature behaviour under the slow-teacher/fast-student split. A geometric planner can bridge two skills; a visual servo can close a local loop. TGL supplies the semantic contract and the feedback structure through which those components contribute to a task, rather than replacing them.
Frequently asked questions
Is TGL a replacement for VLA models?
No. TGL relies on pretrained models for perception, reasoning and control. What it changes is where a newly acquired task capability is stored: in explicit Skill Blocks and memory rather than in the weights, so acquiring one task does not require re-optimising the policy.
Is TGL a world model?
No. TGL does not learn scene dynamics. It stores the semantic effect of a behaviour and recomputes the physical realization from the current observation, which is a different mechanism from predicting a future state.
Does TGL work without an AI Agent?
The architecture is AI Agent-centered by design: the AI Agent selects subgoals, invokes tools, and revises the remaining plan from physical feedback. Specialised robot components still perform the geometry and control, so the method is a division of labour rather than a single model.