---
title: "Teach and Grow: An Agent-Centered Architecture for General Robot Learning"
short_name: "TGL"
expanded_name: "Teach-and-Grow Learning"
canonical: "https://tgl.changnie.top/"
project: "https://tgl.changnie.top/"
paper_page: "https://tgl.changnie.top/paper/"
arxiv: "2608.17209"
arxiv_url: "https://arxiv.org/abs/2608.17209v2"
doi: "10.48550/arXiv.2608.17209"
published: "2026-08-17"
updated: "2026-09-18"
type: "arXiv preprint"
status: "no venue acceptance claimed"
authors:
  - Chang Nie
  - Zhe Liu
  - Hesheng Wang
institution: "Shanghai Jiao Tong University"
language: en
topics:
  - Teach and Grow
  - TGL
  - Teach-and-Grow Learning
  - training-free robot learning
  - robot learning without fine-tuning
  - agent-centered architecture
  - agentic robotics
  - AI agent robot
  - Skill Blocks
  - Skill Library
code: "https://github.com/IRMVLab/TGL"
llms_txt: "https://tgl.changnie.top/llms.txt"
citation: "https://tgl.changnie.top/cite.bib"
---
# Teach and Grow: An Agent-Centered Architecture for General Robot Learning

> Teach and Grow (TGL) is a training-free robot-learning architecture: an AI agent turns sparse teaching into reusable Skill Blocks, acquiring new manipulation tasks with fixed pretrained weights and physical feedback. 99.9% mean success across four LIBERO suites and 92.4% across seven LIBERO-Plus perturbation categories.

Authors: Chang Nie, Zhe Liu, Hesheng Wang — Shanghai Jiao Tong University.
Source: https://tgl.changnie.top/ · Paper: https://tgl.changnie.top/paper/ · Code: https://github.com/IRMVLab/TGL
Facts: https://tgl.changnie.top/project.json · LLM index: https://tgl.changnie.top/llms.txt

---

Teach and Grow (TGL): Training-Free Robot Learning with an AI Agent

**[Skip to content](#main)

AI AGENT-CENTERED ROBOT LEARNING / SHANGHAI JIAO TONG UNIVERSITY

AI Agent robot learning · GPT-6 Astra + Codex
## Teach *and* Grow .

An Agent-Centered Architecture for General Robot Learning

[Chang Nie ↗](https://changnie.top) · Zhe Liu · [Hesheng Wang](mailto:wanghesheng@sjtu.edu.cn)

School of Automation and Intelligent Sensing, Shanghai Jiao Tong University

A few demonstrations become reusable robot skills, with model weights that never change.

[Explore the idea ↓](#overview)[Read the paper · arXiv ↗](https://arxiv.org/abs/2608.17209v2)[Code · GitHub ↗](https://github.com/IRMVLab/TGL)[▷ Watch demonstrations](#demos)

01 / TEACH Sparse demonstrations** Subgoals & shared structure →

02 / ACT + VERIFY **AI Agent + Skill Blocks** Grounded in physical feedback →

03 / GROW **Skills + experience** Retained for the next task

↖ Reusable knowledge returns to the next decision ↵

PRETRAINED WEIGHTS / FIXED EXECUTABLE EXPERIENCE / EVOLVING
[ ↗ Enlarge ](assets/brand/tgl-cover-v4.png)
*New scenes, reusable skills: sparse teaching and AI Agent-guided execution with fixed model weights. Conceptual artwork.*

ABSTRACT / PAPER SUMMARY
## Abstract
**Vision-language-action (VLA) and world-action models typically absorb unfamiliar manipulation tasks through additional robot data collection and policy optimization. This recurring retraining burden slows the acquisition of new behavior. We present Teach-and-Grow Learning (TGL), a training-free architecture that turns a few successful demonstrations into reusable robot skills. Task acquisition requires no gradient updates, fine-tuning, or reinforcement learning: pretrained model weights remain fixed as the robot expands its explicit knowledge. Our implementation uses OpenAI GPT-6 Astra for multimodal reasoning and Codex to connect the AI Agent to robot tools. The AI Agent identifies subgoals shared across demonstrations, expresses them as closed-loop Skill Blocks, and grounds each block in the current scene. Physical feedback guides the next action and any recovery. Verified behaviors enter a persistent Skill Library; Experience Memory records the conditions and repairs that inform later decisions. TGL reaches 99.9% mean success on four LIBERO suites and 92.4% on seven LIBERO-Plus perturbation categories. Controlled studies show that taught blocks persist and improve related-task execution under the same model weights and executors. We further formulate a scaling hypothesis that relates effective reusable experience to falling future-task error and teaching demand. Code and demonstration videos: https://tgl.changnie.top.

Keywords Training-free robot learning · agentic robotics · GPT-6 Astra · Codex · few-shot teaching · vision-language-action models · lifelong learning · skill composition · embodied intelligence

### Where this work sits

Teach and Grow sits inside embodied-intelligence and agentic-robotics research, and targets general robot manipulation. Work in this area is usually organized around vision-language-action (VLA) models, which map perception and language directly to control, and world-action models (WAM), which add learned physical dynamics. Both absorb a new capability by collecting more robot data and optimizing policy parameters. Robot foundation models supply the priors TGL depends on; TGL changes where new task knowledge is stored, and leaves the strength of those priors to the foundation models. Its neighbors therefore include few-shot and sparse-teaching methods, skill-composition and skill-library approaches, lifelong and continual learning, and AI Agent-based robot control.

arXiv preprint · Shanghai Jiao Tong University · 2026

01 / THE PROBLEM
## Where end-to-end robot learning gets expensive.

A robot meets a new container, a changed camera, or a grasp that no longer works. The gap may be small. Paying for it can reach across the whole learning system.

Generalist robot policies have advanced by learning from large collections of observations and actions. Vision-language-action models (VLA) connect semantic understanding to control; world-action models (WAM) add learned physical dynamics. Both routes put new competence into model parameters through training, and that trained prior is what a robot leans on when an unfamiliar task arrives.

Physical interaction data has a different cost structure from text or code. Most of it exists only after a robot or simulator has been run: contact has to be observed, a grasp has to be tried, a collision risk has to be recovered from, and the actual change has to be checked. Object pose, camera geometry, clutter, material, and embodiment also interact, so covering one factor does not cover their combinations.

Repairing a failure through a policy update usually means new data collection, another optimization run, and regression checks on everything the policy already supported. We call this recurring burden the *retraining tax*. It shows up most clearly in the long tail, where a rare contact condition or an unusual object needs one specific lesson rather than another broad round of experience.

The repair is also hard to keep local. If a VLA fails to retain an unfamiliar bowl, corrective data and updated parameters do not leave behind a separately addressable “bowl-retention fix”; the change spreads through the weights, and earlier behavior has to be checked again. A new sensor or gripper repeats the problem at the interface, where observation format, calibration, and action compatibility all need validation.

Long-tail experience is also hard to turn into a reusable correction. A successful trajectory may say nothing about the conditions that mattered after the failed grasp, and when the lesson lives only in parameters, the failure condition, the recovery rule, and the supported scope cannot be inspected or revised one at a time. TGL starts from that question: can a robot keep those objects explicit, and update only the behavior that is affected? RESEARCH QUESTION

Can a robot acquire a new executable capability while keeping its pretrained model weights fixed?

The unit of growth: a verified, reusable behavior.

TGL explores a different update mechanism. A few successful demonstrations give an initial task structure; the robot turns that structure into explicit skills, tests them by execution, and keeps what worked together with the experience that explains it. The pretrained stack still supplies the priors. What changes is where new task knowledge lives. [ ↗ Enlarge ](assets/figures/fig1.webp)
*The paper’s conceptual comparison: policy optimization on one side, explicit skill acquisition and verification on the other.*

02 / HOW WE ARRIVED HERE
## From an AI Agent that can plan to a robot that can grow.

The architecture grew out of a practical question: what is still missing when a multimodal AI Agent understands a task but cannot reliably carry it out?

### From LLMs and ChatGPT to tool-using AI Agents.

Large language models (LLMs), familiar through assistants such as ChatGPT and Claude, brought language understanding and reasoning into an interactive interface. Coding AI Agents such as [Codex](https://developers.openai.com/codex/) and [Claude Code](https://code.claude.com/docs/en/overview) add a working pattern that transfers well to a robot: inspect the current state, choose a tool, act, look at the result, and revise the plan.

TGL moves that loop into physical manipulation. Camera images stand in for workspace state, perception and grasping and motion tools carry out the operations, and new observations report what the action changed. From that evidence the AI Agent chooses the next subgoal and the next tool. Reasoning organizes the task; specialist robot components supply the physical competence.

The paper’s implementation uses OpenAI GPT-6 Astra for multimodal reasoning and Codex as the AI Agent interface . ChatGPT and Claude appear here only as the broader LLM context, and Claude Code illustrates the coding-AI Agent pattern. That connection raised the question behind the rest of the work: how much of the pattern survives when a wrong action changes the physical world?

01

### Direct visual planning

We began by asking a multimodal AI Agent to plan from robot-view images. It could interpret the goal, but the physical gaps remained: metric depth, collision-free motion, contact, high-rate control. That first attempt separated semantic competence from the robot-native competence a plan needs before it can run at all.

02

### Tools and a return path for feedback

We then connected perception, grasping, planning, and control tools, and returned a fresh observation after each short execution. When the scene disagreed with the plan, the AI Agent could revise its route. Structured memory made one trial worth more than its episode by keeping its conditions, outcome, and repair.

03

### Sparse teaching as a productive starting point

Tool use made zero-shot AI Agent control feasible, yet learning one behavior by repeated trial still consumed interaction and model calls. Successful teaching supplies a subgoal sequence early, so exploration can concentrate on the variation, correction, and missing transitions the demonstrations do not cover. This became the Teach-and-Grow route.

教学相长

Teaching gives exploration a starting structure. Execution exposes what the teaching left unresolved. Those gaps make the next teaching intervention more focused.

### Why not simply let the AI Agent drive the robot and finish the task on its own?

An AI Agent driving the robot through perception, grasping, planning and control tools can complete a manipulation task with no teaching at all; we ran that experiment, and this work began there. What it does not do is accumulate. Every run reasons from scratch, the route that worked and the grasp that succeeded disappear when the episode ends, and the same task is solved again at the same cost. TGL keeps the AI Agent loop and changes where the result is stored. A successful task becomes explicit Skill Blocks with a stated scope and an outcome test, and the conditions and repairs go into Experience Memory. From then on the AI Agent retrieves instead of re-deriving, converges faster because the subgoal structure is already settled, reuses verified behavior across tasks, and carries the experience forward. That is the difference between an AI Agent that can solve a task and a system that keeps what it solved.

### Why teach at all? Would autonomous exploration not be better?

Autonomous exploration is possible and zero-shot AI Agent control is real; we ran it. It is a poor default for two reasons. It converges slowly, because learning one physical behavior by trial and error costs many robot interactions and many sequential model calls. And on real hardware it cannot be left unbounded, because an exploratory action is a physical action that can damage the object, the gripper or the scene. Teaching changes the starting point rather than the goal: a few successful demonstrations supply the subgoal order and the conditions worth checking, so exploration concentrates on the variation, correction and recovery the demonstrations did not cover. Teaching is an accelerator, not a precondition, and its sources are open. Robot trajectories, simulation, human video and a written procedure all count, because anything that shows how the task is completed and what it accomplishes can seed the structure.

Teaching reduced the search for a useful starting plan. The next question was what to keep from it: the record of one successful episode, or a behavior that could be grounded and executed in another scene?

03 / THE CENTRAL IDEA
## Make executable capability the object of learning.

A demonstration holds two different things: a strategy worth keeping, and physical details that belong to one scene. TGL separates them.

Consider placing a bowl on a plate. Across demonstrations the hand may approach from different directions and follow different paths. The shared structure is steadier: acquire the requested bowl, establish that it is held, move toward the target relation, and release it onto the plate.

TGL represents one meaningful change as a Skill Block . The block keeps the goal, the reusable strategy, the conditions under which it applies, and the test of its outcome. The current observation supplies the actual object pose, grasp, path, and control target. Reuse therefore preserves the intended effect while the physical realization stays free to change.

Training-free has a precise meaning here: acquiring the incoming task invokes no gradient update, no fine-tuning, and no reinforcement-learning stage. The AI Agent and its specialist models may already be pretrained. What changes during acquisition is the explicit skill and memory state.

θ

Pretrained model weights** Fixed during task acquisition FIXED

B

**Skill Library** Executable behaviors, scopes, and contracts GROWS

M

**Experience Memory** Conditions, outcomes, diagnoses, and repairs GROWS

A task leaves the robot with something it can retrieve, inspect, and execute again.

### Does the operating system replace VLA and world-action models?

No, and it is not designed to. TGL is a new AI Agent-driven general-purpose robot operating system, and the models it organizes remain the strongest way to turn perception into fast continuous control. In that arrangement a trained policy is a component the system calls: a policy that reliably picks one category of object or opens one kind of drawer can be registered as the executor of a Skill Block, and the AI Agent selects it, grounds it in the current scene, checks its outcome and composes it with other blocks. Where a VLA is strong, TGL uses it directly and benefits from every improvement to it; where a task is new, rare or outside the policy distribution, the AI Agent acquires the missing structure explicitly instead of waiting for the next training round. The two are complements at different levels of the stack, not rivals in the same slot.

RETAIN
### Semantic effects & strategy

What should change, in which order, under which conditions. ×

RECOMPUTE
### Current physical realization

Object bindings, grasp geometry, collision-free motion, and control.

The separation defines what has to be learned. Making it work on a robot adds a second connection: every retained subgoal has to lead to an action in the current scene, and every action has to return evidence that can change the next decision.

04 / INSIDE THE ARCHITECTURE
## Subgoal reasoning on top, geometric control underneath.

The AI Agent organizes the task at meaningful state transitions. Inside each block, specialized executors handle geometry and continuous control. [ ↗ Enlarge ](assets/figures/fig2.webp)
*One task moves through teaching, composition, execution, verification, and persistent storage. The resulting library and memory become the starting point for the next task.*

A

### Read demonstrations as evidence about change

Demonstrations are cut at meaningful events, such as acquiring an object or establishing a containment relation. Segments align by what they accomplish, even when their timing and motion differ. The shared structure becomes a candidate strategy, and the variation across demonstrations decides how broadly it can be claimed: an example from one object stays narrow until more evidence supports a wider scope.

B

### Give each skill a contract with the scene

A block answers practical questions. What effect is intended? When does it apply? What evidence does it need? Which executors can realize it? How is success observed? What recovery is allowed? During acquisition, a closed gripper alone does not establish that the object is held; verification has to refer to the intended physical effect.

C

### Revise the remaining plan after the robot acts

The working plan is an ordered composition of blocks, and its remainder can change. A passed effect opens the next stage; a failed or inconclusive effect can call for another observation, a different executor, or a revised route. This is where reasoning meets physical feedback: what happens next depends on what actually happened.

D

### Validate before making a candidate reusable

A promising candidate is evaluated beyond its teaching demonstrations. Supported scope, executor compatibility, outcome test, and recovery are checked before it enters the library; weak candidates get narrowed or repaired. Library growth is therefore a decision about one behavior and the conditions under which it works, not an automatic consequence of every episode. [ ↗ Enlarge ](assets/figures/fig3.webp)
*The contract connects a semantic subgoal to grounded execution and observable evidence.*

A WORKED CONCEPTUAL EXAMPLE
### “Put the bowl on the plate.”

Follow which information is retained and which decisions are made again.

01 Read the teaching 02 Ground the block 03 Read the outcome 04 Keep the lesson

Read the teaching
#### Acquire → transport → release

The demonstrations supply the subgoals and their order. They also show which conditions are worth checking, such as holding the bowl before transport starts. Their original pixel coordinates and exact trajectory are not part of the reusable strategy.

Ground the block
#### Current bowl → feasible grasp

The block binds the requested object to the current observation, and a compatible grasping and motion backend computes the action. The contact point can differ from the teacher’s while the state the block is after stays the same.

Read the outcome
#### Effect → continue / reobserve / repair

After execution, a verifier checks whether the bowl is actually held. A weak or failed grasp changes the next decision: the AI Agent can look again, or revise the remaining route before it attempts placement.

Keep the lesson
#### Validated behavior + useful context

The executable strategy and its scope enter the Skill Library once validated. The conditions of the attempt, its outcome, and a useful repair enter Experience Memory. A later task can retrieve both.

### Why call external tools such as object detection if the AI Agent can already see?

Because the AI Agent’s own visual judgment is not precise enough to manipulate with. In our experiments it handled simple tasks from its own image understanding alone: which object is meant, a rough order of operations, whether a scene looks like the goal. Measurement is where it becomes unreliable. A 6-DoF pose, a centimetre-scale clearance, the boundary of an occluded object, whether the gripper is really holding something, whether a path would collide: those need metric answers that hold from one frame to the next, and a general-purpose model reading a camera image gives approximate ones. Detection and segmentation establish object identity and boundary, RGB-D geometry supplies depth and pose, grasp and motion planners produce collision-free reachable actions, and controllers hold the loop at the required rate. The AI Agent is then free to organize the task, choose the next subgoal, read the outcome and decide what comes next. Because each tool sits inside a block, a better detector can also be swapped in and retested without disturbing the reasoning or the library.

In the paper’s current system, OpenAI GPT-6 Astra and Codex provide task-level reasoning and tool interaction. Detection, segmentation, RGB-D geometry, Contact-GraspNet, MPLib, and controllers supply the physical grounding and execution.

05 / WHAT THE ROBOT KEEPS
## A skill stores behavior. Memory stores its lessons.

Solving one episode and acquiring a lasting capability are different events. TGL keeps the two objects separate.

B / EXECUTABLE
### Skill Library

The library holds behaviors that can be selected and run: their subgoals, supported conditions, grounding rules, compatible tools, and verification logic. A text description of “how to grasp” is not enough. The block has to connect that intent to an executor and to a test of its effect.

M / CONTEXTUAL
### Experience Memory

Memory keeps the context of use: the task, the blocks that were selected, the observations, the outcome, the diagnosis, the repair. A failed attempt may reveal an unsuitable grasp family or an ambiguous observation, and keeping that explanation can guide the next selection without turning every episode into a new executable block.

The separation also keeps the system inspectable. A person can narrow an overgeneralized scope, change a recovery rule, or mark a tool version as incompatible. Keeping a file is not the same as keeping a behavior: the right block still has to be retrieved, grounded in current sensing, and executed successfully, and those remain separate questions for a robot that learns over years.

The paired videos below show the loop directly. Start from the teacher’s task structure, then watch how TGL reaches the same outcome through a different grasp, or what it does when the first attempt fails to secure the object.

06 / BEHAVIOR IN CONTEXT
## Learn the task. Adapt the execution.

Five paired LIBERO demonstrations put the teacher’s example next to Teach and Grow on the same task. Three suite families appear here: Object tasks (place a named object in a basket), Spatial tasks (place one object in relation to another), and Goal tasks (bring about a state, such as turning on the stove).

All five Object Spatial Goal Teacher ← → TGL · videos play on request

Qualitative simulation examples. Clips have different durations; joint playback starts them together without aligning individual actions. Playback speed refers to the supplied videos, not measured robot latency.

OBJECT /01
### Same goal. A different grasp.

Place the soup can in the basket Adaptation

Teacher demonstration Your browser does not support embedded video. [MP4](assets/videos/libero_object_task00_teacher.mp4)

Teach and Grow · ours Your browser does not support embedded video. [MP4](assets/videos/libero_object_task00_system.mp4)

▶ Play both ↺ Restart Speed 0.5× 1× 2×

The teacher supplies the task structure: acquire the can, carry it, release it in the basket. TGL takes a different approach direction and a different grasp, and still reaches the same outcome. The strategy carried over; the motion did not.

**Watch for:** Compare the approach direction and gripper orientation before the lift.

SPATIAL /02
### A new contact point. The same subgoal.

Place the bowl on the plate Scene grounding

Teacher demonstration Your browser does not support embedded video. [MP4](assets/videos/libero_spatial_task00_teacher.mp4)

Teach and Grow · ours Your browser does not support embedded video. [MP4](assets/videos/libero_spatial_task00_system.mp4)

▶ Play both ↺ Restart Speed 0.5× 1× 2×

TGL grasps a different part of the bowl rim than the teacher does. The subgoal is unchanged; the contact point is recomputed from the scene in front of the robot.

**Watch for:** Follow where the fingers contact the rim in each rollout.

SPATIAL /03
### Miss. Observe. Try again.

Move the bowl from the cabinet to the plate Feedback & recovery

Teacher demonstration Your browser does not support embedded video. [MP4](assets/videos/libero_spatial_task09_teacher.mp4)

Teach and Grow · ours Your browser does not support embedded video. [MP4](assets/videos/libero_spatial_task09_system.mp4)

▶ Play both ↺ Restart Speed 0.5× 1× 2×

The first grasp does not secure the bowl. In this system rollout, TGL tries again and completes the placement: recovery driven by what the robot observed after the failed lift.

**Watch for:** Watch the first lift attempt, then the renewed approach and successful transfer.

OBJECT /04
### From a taught sequence to execution.

Place the dressing bottle in the basket Skill execution

Teacher demonstration Your browser does not support embedded video. [MP4](assets/videos/libero_object_task02_teacher.mp4)

Teach and Grow · ours Your browser does not support embedded video. [MP4](assets/videos/libero_object_task02_system.mp4)

▶ Play both ↺ Restart Speed 0.5× 1× 2×

The teacher and TGL both complete the bottle transfer. On a second object the same three stages hold: acquire, transport, release.

**Watch for:** Track the bottle from the initial grasp to release above the basket.

GOAL /05
### Turn a goal into a physical change.

Turn on the stove Articulated control

Teacher demonstration Your browser does not support embedded video. [MP4](assets/videos/libero_goal_task07_teacher.mp4)

Teach and Grow · ours Your browser does not support embedded video. [MP4](assets/videos/libero_goal_task07_system.mp4)

▶ Play both ↺ Restart Speed 0.5× 1× 2×

Here the goal is not a change of location. The robot has to reach the stove knob and turn it, and both clips follow the approach and the rotation itself.

**Watch for:** Focus on the gripper–knob contact and the resulting rotation.

07 / WHAT WE INVESTIGATED
## Four studies of how the mechanism behaves.

The videos show single scenes. The studies look behind them: where a skill comes from, what survives a task, and what changes the next execution.

### Demonstration → semantic effects

Ten saved demonstrations from two related LIBERO-Object tasks are split into 40 predicted stages. Ordered role/type accuracy is 1.000, and all 20 observable acquire/release effects are confirmed. Frame boundaries are much weaker: exact-frame F1 is 0.10, rising to 0.90 within two sampled frames. The study is deliberately narrow, using deterministic visual cues and typed rules to isolate the representation and verification mechanism, and it separates two things cleanly. The stage sequence is recovered; the exact frame is not.

### Candidate → persistent Skill Block

Three teacher trajectories produce two blocks: acquire the requested object, release it in the required relation. On separated states the pair succeeds 3/3, and it still succeeds 3/3 after save-and-reload. On farther states 6–8, the route stops at the first missing semantic effect, which is what keeps a local mistake from propagating through the task.

### Physical feedback → a revised decision

Online traces inspect two decisions: replanning the remaining route in a bowl-placement task, and requesting another observation when the drawer-opening verification comes back inconclusive. A separate failure cohort locates where the remaining attempts broke: two before motion planning, four in path consistency or calibration, two at gripper closure.

### Library growth → related-task reuse

One controlled pilot holds executor, seeds, and budget fixed, and adds only the learned acquisition and release blocks. Related-task success changes only once those two blocks enter the library, with the executor and the evaluation conditions untouched. The sample is small, so the paper reports it as a directional observation rather than a benchmark claim.

Experimental protocols, quantitative results, and comparison tables are available in the paper. [Read the evaluation ↗](https://arxiv.org/pdf/2608.17209v2#page=4)

08 / THE LARGER RESEARCH DIRECTION
## Scale the experience a robot can actually reuse.

These studies follow the path from a taught task to a reusable skill. The next question concerns a longer timescale: what happens as that process repeats throughout a robot’s working life?

More stored trajectories do not necessarily mean more capability. Experience is useful only if it can still be retrieved, grounded, and composed for a new scene. The Teach-and-Grow hypothesis relates this effective reusable experience to declining future-task error and teaching demand. It is a hypothesis about lifelong capability growth, and the present experiments do not yet fit a universal scaling law. [ ↗ Enlarge ](assets/figures/fig4.webp)
*Conceptual predictions: more reusable experience can reduce future error and the teaching required for related tasks. The curves are schematic.*

### Acquire unfamiliar behavior deliberately; execute familiar behavior efficiently.

Agentic acquisition uses repeated observation, reasoning, and tool calls. That deliberation is useful at the frontier of the robot’s knowledge, while a learned policy handles mature behavior under the slow-teacher/fast-student split.

### The AI Agent is slower than a policy. How is that handled?

On unfamiliar tasks it is slow, and that trade is deliberate: deliberation sits at subgoal boundaries while the executors below run continuous control at full rate. Three things follow. First, the cost is falling on its own, because each generation of reasoning models is faster and cheaper at the same capability, and prompt caching, tool-call batching and stronger multimodal perception keep reducing the number of sequential calls a task needs, so an architecture that reasons at semantic boundaries benefits from that trend directly. Second, slow acquisition is what makes TGL a strong data source. A general method that can acquire an unfamiliar manipulation task produces verified trajectories along the way: real observations, real actions and real outcomes, already checked against an explicit success criterion. That is exactly the supervision a small fast policy, a VLA or a WAM needs, and it covers the long-tail conditions that are hardest to collect by hand, so the system serves as a data generator for the fast models it later calls. Third, the two layers occupy different positions rather than competing: the agentic layer is strong and slow and handles new tasks, rare conditions and data collection, while the fast layer is narrow and quick and runs mature behavior at policy speed. Data flows from the slow side to the fast side, and when a fast policy meets something outside its competence, control returns to the AI Agent, which diagnoses the gap and grows the library.

### VLA, WAM, and classical robotics remain part of the system.

A learned policy can implement a block, execute a familiar composition, or become a future student. A geometric planner can bridge two skills; a visual servo can close a local loop; tactile sensing can strengthen a contact check. TGL supplies the semantic contract and feedback structure through which these components contribute to a task.

The same perspective extends to teaching sources. Robot trajectories, simulation, human video, and written procedures offer different kinds of evidence. A manual may reveal the order of operations while leaving the grasp unresolved. The architecture preserves that distinction: semantic knowledge guides acquisition, and robot-specific grounding and validation determine what can actually run.

### Other methods look similar. What is new here?

Several recent systems work on neighbouring pieces, and the paper cites them: LRLL, ASPIRE, SkillMemo, SCE and PACTS study lifelong skill acquisition, agentic discovery, memory and compositional reuse, while PhyAgentOS, AEROS and RoboBridge build robot operating layers. Individual ingredients — skill libraries, agentic tool use, episodic memory, demonstration decomposition — are not new. TGL is the first system to propose the method as a whole: an AI Agent-driven general-purpose robot operating system that connects sparse teaching, explicit closed-loop Skill Blocks carrying a scope and an outcome test, weight-frozen execution, physical feedback and recomposition, structured failure memory, and persistent growth into one single learning cycle. Neighbouring methods cover parts of that cycle, so the contribution here is the cycle itself and the interfaces between its parts. That whole-system view is what makes faster convergence, verified reuse, experience retained across tasks and embodiments, and the role of data source for fast policies available together rather than one at a time. [ ↗ Enlarge ](assets/figures/fig6.webp)
*A whole-system view. Dashed paths mark trajectories routed to fast-policy training and sharing across a robot fleet.*

### Why local skill updates may change acquisition cost

A local skill addition can avoid reopening every part of the learning system, provided grounding, validation, compatibility, and retrieval remain manageable. Those conditions matter: unrestricted pairwise compatibility checks can make library growth expensive. The paper analyzes these different cost regimes rather than treating low-cost growth as automatic. [ ↗ Enlarge ](assets/figures/fig5.webp)
*Analytic cost regimes under different assumptions about coverage and local compatibility.*

09 / PAPER & RESOURCES
## Read the full paper.

The paper covers the formulation, the Skill Block contract, the LIBERO and LIBERO-Plus evaluations, the controlled studies, and the scaling hypothesis, with appendices on cost regimes and experimental details. [Read the paper · arXiv ↗](https://arxiv.org/abs/2608.17209v2)

[arXiv:2608.17209 ↗](https://arxiv.org/abs/2608.17209v2)[Method code · IRMVLab/TGL ↗](https://github.com/IRMVLab/TGL)[10 demonstration videos ↗](#demos)[Chang Nie · Personal homepage ↗](https://changnie.top)[Contact the authors ↗](mailto:changniep@gmail.com)

The method implementation is maintained in IRMVLab/TGL. Installation and execution instructions are in its README.

### BibTeX
Copy citation

`@misc{nie2026teachandgrow,
title = {Teach and Grow: An Agent-Centered Architecture for General Robot Learning},
author = {Nie, Chang and Liu, Zhe and Wang, Hesheng},
year = {2026},
eprint = {2608.17209},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
doi = {10.48550/arXiv.2608.17209},
url = {https://arxiv.org/abs/2608.17209v2}
}`

GLOSSARY / KEY TERMS
## Key terms

Teach-and-Grow Learning (TGL) A training-free robot-learning architecture that converts a few successful demonstrations into reusable, verifiable skills while pretrained model weights remain fixed. New task knowledge lives in the skill and memory stores, not in the weights.

Training-free robot learning Acquiring a new robot capability without gradient updates, fine-tuning, or reinforcement learning. Pretrained weights stay fixed; new task knowledge lives in explicit skill and memory stores.

Skill Block The unit of reusable robot behavior in TGL: a goal, a reusable strategy, supported conditions, compatible executors, and an outcome test. The semantic effect is retained; the physical realization is recomputed from the current scene. Success is decided by that effect, so a closed gripper does not by itself pass an acquisition block.

Skill Library The persistent store of validated Skill Blocks, including their scopes, contracts, and executor compatibility. It grows after validation, not after every episode, so a stored file is not the same as a retained behavior.

Experience Memory The contextual store recording the task, selected blocks, observations, outcome, diagnosis, and repair of an attempt, so later decisions can reuse the conditions as well as the behavior.

Retraining tax The recurring cost of repairing robot behavior through policy updates: new data collection, optimization, and regression checking against previously supported behavior.

FAQ / QUESTIONS
## Questions about the method

### What is Teach and Grow?

Teach and Grow (TGL) is a training-free architecture for general robot learning. A pretrained AI Agent turns a few successful demonstrations into explicit, reusable skills, so the robot acquires new manipulation tasks while its pretrained model weights stay fixed.

### What does “training-free” mean here exactly?

It means acquiring the incoming task invokes no gradient update, no fine-tuning, and no reinforcement-learning stage. The AI Agent and its specialist models may already be pretrained; what changes during task acquisition is the explicit skill and memory state, not the weights.

### What is the retraining tax?

It is the recurring cost of repairing robot behavior through a policy update: new data collection, another optimization run, and regression checks on everything the policy already supported. It is called a tax because the cost returns with every new task, sensor, or gripper, and grows with the amount of behavior that already works.

### How is this different from training a VLA or world-action model?

VLA and world-action models absorb a new behavior by collecting more robot data and optimizing policy parameters. TGL repairs and extends behavior through explicit Skill Blocks instead. A learned policy can still take part: it may act as the executor inside a block, or later serve as a student of verified trajectories.

### What is a Skill Block?

A Skill Block is the unit of reusable behavior: a goal, a reusable strategy, supported conditions, compatible executors, and an outcome test. The semantic effect is what gets retained, while the physical realization (object bindings, grasp geometry, collision-free motion) is recomputed from the current scene. An acquisition block, for instance, is never satisfied by a closed gripper alone.

### What robot and AI Agent does the implementation use?

The implementation uses OpenAI GPT-6 Astra for multimodal reasoning and Codex to connect the AI Agent to robot tools. Detection, segmentation, RGB-D geometry, Contact-GraspNet, MPLib and controllers supply the physical grounding and execution, evaluated in the LIBERO simulation suite.

### Why not just let GPT, Codex or Claude Code drive the robot and finish the task in one go?

Because finishing a task once and building a system that keeps getting better at it are different things. We started from exactly that experiment: an AI Agent driving the robot through perception, grasping, planning and control tools completed manipulation tasks with no teaching at all. What did not happen is accumulation. Every run reasoned from scratch, the successful route and the grasp that worked disappeared when the episode ended, and a mature behavior was never cheaper the second time. TGL keeps the same AI Agent loop and changes where the result is stored. A task that succeeds is split into explicit Skill Blocks with a stated scope and an outcome test, and the conditions and repairs go into Experience Memory. From then on the AI Agent retrieves an existing skill instead of re-deriving it, converges faster because the subgoal structure is already settled, reuses verified behavior across tasks, and carries the accumulated experience into later work. An AI Agent can solve a task on its own; the architecture is what turns that one-off success into a system that grows, rather than a run that is thrown away after use.

### Why few-shot teaching? Isn't autonomous exploration enough?

Autonomous exploration works, and we ran zero-shot AI Agent control, so this is a design choice rather than a limitation. It is a poor default for two reasons. Convergence is slow, because learning one physical behavior by trial and error consumes many robot interactions and many sequential model calls, as every attempt needs a fresh decision. And on real hardware it is unsafe to leave unbounded, since an exploratory action is a physical action and a wrong one can damage the object, the gripper or the scene. Few-shot teaching changes the starting point rather than the goal: a few successful demonstrations supply the subgoal order and the conditions worth checking, so exploration can concentrate on the variation, correction and recovery the demonstrations did not cover. Teaching is an accelerator, not a precondition. The source of a demonstration is open as well. Robot trajectories, simulation, human video and a written procedure all count, because anything that shows how the task is completed and what it accomplishes can provide the initial structure. A manual may reveal the order of operations while leaving the grasp unresolved, and the architecture keeps that distinction: semantic knowledge guides acquisition, while robot-specific grounding and validation decide what can actually run.

### Why call external tools such as object detection if the AI Agent can already see?

Because the AI Agent's own visual judgment is not precise enough to manipulate with. In our experiments it handled simple tasks from its own image understanding alone: recognizing which object is meant, choosing a rough order of operations, judging whether a scene looks like the goal. Measurement is where it becomes unreliable. A 6-DoF pose, a centimetre-scale clearance, the boundary of an occluded object, whether the gripper is actually holding something, whether a path would collide: these need metric answers that hold from one frame to the next, and a general-purpose model reading a camera image gives approximate ones. Dedicated tools close that gap. Detection and segmentation establish object identity and boundary, RGB-D geometry supplies depth and pose, grasp and motion planners produce collision-free reachable actions, and controllers hold the loop at the required rate. The AI Agent is then free to do what it is good at: organizing the task, choosing the next subgoal, reading the outcome, deciding what comes next. The split also keeps the system maintainable, because a detector is a replaceable component inside a Skill Block that can be swapped and retested without touching the reasoning, the library or the rest of the pipeline.

### How does TGL relate to VLA and world-action models?

TGL is a new AI Agent-driven general-purpose robot operating system, and it is designed to work with VLA and world-action models rather than replace them. The two sit at different levels. A VLA or WAM remains the strongest available way to turn perception into fast continuous control; TGL supplies the operating layer above it, which decides what subgoal comes next, which capability applies in this scene, what actually happened, and what should be kept. In that arrangement a learned policy is a component the system calls. A policy trained for a bounded task, such as picking one category of object or opening one kind of drawer, can be registered as the executor of a Skill Block: the AI Agent selects it, grounds it in the current scene, checks its outcome and composes it with other blocks, and a mature block can run entirely inside the fast policy once the behavior is stable. The relationship is therefore composition, not competition. Where a VLA is strong, TGL uses it directly and benefits from every improvement to it; where a task is new, rare or outside the policy's distribution, the AI Agent acquires the missing structure explicitly instead of waiting for the next training round.

### Isn't an AI Agent-driven system too slow?

On unfamiliar tasks it is slow, and that trade is deliberate: deliberation is placed at subgoal boundaries while the executors below run continuous control at full rate. Three things follow. First, the cost is falling on its own, because each generation of reasoning models is faster and cheaper at the same capability, and prompt caching, tool-call batching and stronger multimodal perception keep reducing the number of sequential calls a task needs, so an architecture that puts reasoning at semantic boundaries benefits from that trend directly. Second, slow acquisition is what makes TGL a strong data source. A general method that can acquire an unfamiliar manipulation task produces, along the way, verified trajectories: real observations, real actions and real outcomes, already checked against an explicit success criterion. That is exactly the supervision a small fast policy, a VLA or a WAM needs, and it covers precisely the long-tail conditions that are hardest to collect by hand, so the system serves as a data generator for the fast models it later calls. Third, the two layers occupy different positions rather than competing: the agentic layer is strong and slow and handles new tasks, rare conditions and data collection, while the fast layer is narrow and quick and runs mature behavior at policy speed. Data flows from the slow side to the fast side, and when a fast policy meets something outside its competence, control returns to the AI Agent, which diagnoses the gap and grows the library.

### Other methods look similar. What is new here?

Several recent systems work on neighbouring pieces, and the paper cites them: LRLL, ASPIRE, SkillMemo, SCE and PACTS study lifelong skill acquisition, agentic discovery, memory and compositional reuse, while PhyAgentOS, AEROS and RoboBridge build robot operating layers. Individual ingredients such as skill libraries, agentic tool use, episodic memory and demonstration decomposition are not new, and TGL does not claim them. TGL is the first system to propose the method as a whole: an AI Agent-driven general-purpose robot operating system that connects sparse teaching, explicit closed-loop Skill Blocks carrying a scope and an outcome test, weight-frozen execution, physical feedback and recomposition, structured failure memory, and persistent growth into one single learning cycle. Each neighbouring method covers part of that cycle, so the contribution here is the cycle itself and the interfaces between its parts. That whole-system view is what makes the practical properties available, namely converging faster on a new task, reusing a verified behavior instead of re-deriving it, keeping experience across tasks and embodiments, and serving as a data source for fast policies.

---

## Citation

```bibtex
@misc{nie2026teachandgrow,
  title = {Teach and Grow: An Agent-Centered Architecture for General Robot Learning},
  author = {Nie, Chang and Liu, Zhe and Wang, Hesheng},
  year = {2026},
  eprint = {2608.17209},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  doi = {10.48550/arXiv.2608.17209},
  url = {https://arxiv.org/abs/2608.17209v2}
}
```
