FAQ

What is Teach and Grow?

Teach and Grow (TGL) is a training-free architecture for general robot learning. A pretrained AI Agent turns a few successful demonstrations into explicit, reusable skills, so the robot acquires new manipulation tasks while its pretrained model weights stay fixed.

What does “training-free” mean here exactly?

It means acquiring the incoming task invokes no gradient update, no fine-tuning, and no reinforcement-learning stage. The AI Agent and its specialist models may already be pretrained; what changes during task acquisition is the explicit skill and memory state, not the weights.

What is the retraining tax?

It is the recurring cost of repairing robot behavior through a policy update: new data collection, another optimization run, and regression checks on everything the policy already supported. It is called a tax because the cost returns with every new task, sensor, or gripper, and grows with the amount of behavior that already works.

How is this different from training a VLA or world-action model?

VLA and world-action models absorb a new behavior by collecting more robot data and optimizing policy parameters. TGL repairs and extends behavior through explicit Skill Blocks instead. A learned policy can still take part: it may act as the executor inside a block, or later serve as a student of verified trajectories.

What is a Skill Block?

A Skill Block is the unit of reusable behavior: a goal, a reusable strategy, supported conditions, compatible executors, and an outcome test. The semantic effect is what gets retained, while the physical realization (object bindings, grasp geometry, collision-free motion) is recomputed from the current scene. An acquisition block, for instance, is never satisfied by a closed gripper alone.

What robot and AI Agent does the implementation use?

The implementation uses OpenAI GPT-6 Astra for multimodal reasoning and Codex to connect the AI Agent to robot tools. Detection, segmentation, RGB-D geometry, Contact-GraspNet, MPLib and controllers supply the physical grounding and execution, evaluated in the LIBERO simulation suite.

Why not just let GPT, Codex or Claude Code drive the robot and finish the task in one go?

Because finishing a task once and building a system that keeps getting better at it are different things. We started from exactly that experiment: an AI Agent driving the robot through perception, grasping, planning and control tools completed manipulation tasks with no teaching at all. What did not happen is accumulation. Every run reasoned from scratch, the successful route and the grasp that worked disappeared when the episode ended, and a mature behavior was never cheaper the second time. TGL keeps the same AI Agent loop and changes where the result is stored. A task that succeeds is split into explicit Skill Blocks with a stated scope and an outcome test, and the conditions and repairs go into Experience Memory. From then on the AI Agent retrieves an existing skill instead of re-deriving it, converges faster because the subgoal structure is already settled, reuses verified behavior across tasks, and carries the accumulated experience into later work. An AI Agent can solve a task on its own; the architecture is what turns that one-off success into a system that grows, rather than a run that is thrown away after use.

Why few-shot teaching? Isn't autonomous exploration enough?

Autonomous exploration works, and we ran zero-shot AI Agent control, so this is a design choice rather than a limitation. It is a poor default for two reasons. Convergence is slow, because learning one physical behavior by trial and error consumes many robot interactions and many sequential model calls, as every attempt needs a fresh decision. And on real hardware it is unsafe to leave unbounded, since an exploratory action is a physical action and a wrong one can damage the object, the gripper or the scene. Few-shot teaching changes the starting point rather than the goal: a few successful demonstrations supply the subgoal order and the conditions worth checking, so exploration can concentrate on the variation, correction and recovery the demonstrations did not cover. Teaching is an accelerator, not a precondition. The source of a demonstration is open as well. Robot trajectories, simulation, human video and a written procedure all count, because anything that shows how the task is completed and what it accomplishes can provide the initial structure. A manual may reveal the order of operations while leaving the grasp unresolved, and the architecture keeps that distinction: semantic knowledge guides acquisition, while robot-specific grounding and validation decide what can actually run.

Why call external tools such as object detection if the AI Agent can already see?

Because the AI Agent's own visual judgment is not precise enough to manipulate with. In our experiments it handled simple tasks from its own image understanding alone: recognizing which object is meant, choosing a rough order of operations, judging whether a scene looks like the goal. Measurement is where it becomes unreliable. A 6-DoF pose, a centimetre-scale clearance, the boundary of an occluded object, whether the gripper is actually holding something, whether a path would collide: these need metric answers that hold from one frame to the next, and a general-purpose model reading a camera image gives approximate ones. Dedicated tools close that gap. Detection and segmentation establish object identity and boundary, RGB-D geometry supplies depth and pose, grasp and motion planners produce collision-free reachable actions, and controllers hold the loop at the required rate. The AI Agent is then free to do what it is good at: organizing the task, choosing the next subgoal, reading the outcome, deciding what comes next. The split also keeps the system maintainable, because a detector is a replaceable component inside a Skill Block that can be swapped and retested without touching the reasoning, the library or the rest of the pipeline.

How does TGL relate to VLA and world-action models?

TGL is a new AI Agent-driven general-purpose robot operating system, and it is designed to work with VLA and world-action models rather than replace them. The two sit at different levels. A VLA or WAM remains the strongest available way to turn perception into fast continuous control; TGL supplies the operating layer above it, which decides what subgoal comes next, which capability applies in this scene, what actually happened, and what should be kept. In that arrangement a learned policy is a component the system calls. A policy trained for a bounded task, such as picking one category of object or opening one kind of drawer, can be registered as the executor of a Skill Block: the AI Agent selects it, grounds it in the current scene, checks its outcome and composes it with other blocks, and a mature block can run entirely inside the fast policy once the behavior is stable. The relationship is therefore composition, not competition. Where a VLA is strong, TGL uses it directly and benefits from every improvement to it; where a task is new, rare or outside the policy's distribution, the AI Agent acquires the missing structure explicitly instead of waiting for the next training round.

Isn't an AI Agent-driven system too slow?

On unfamiliar tasks it is slow, and that trade is deliberate: deliberation is placed at subgoal boundaries while the executors below run continuous control at full rate. Three things follow. First, the cost is falling on its own, because each generation of reasoning models is faster and cheaper at the same capability, and prompt caching, tool-call batching and stronger multimodal perception keep reducing the number of sequential calls a task needs, so an architecture that puts reasoning at semantic boundaries benefits from that trend directly. Second, slow acquisition is what makes TGL a strong data source. A general method that can acquire an unfamiliar manipulation task produces, along the way, verified trajectories: real observations, real actions and real outcomes, already checked against an explicit success criterion. That is exactly the supervision a small fast policy, a VLA or a WAM needs, and it covers precisely the long-tail conditions that are hardest to collect by hand, so the system serves as a data generator for the fast models it later calls. Third, the two layers occupy different positions rather than competing: the agentic layer is strong and slow and handles new tasks, rare conditions and data collection, while the fast layer is narrow and quick and runs mature behavior at policy speed. Data flows from the slow side to the fast side, and when a fast policy meets something outside its competence, control returns to the AI Agent, which diagnoses the gap and grows the library.

Other methods look similar. What is new here?

Several recent systems work on neighbouring pieces, and the paper cites them: LRLL, ASPIRE, SkillMemo, SCE and PACTS study lifelong skill acquisition, agentic discovery, memory and compositional reuse, while PhyAgentOS, AEROS and RoboBridge build robot operating layers. Individual ingredients such as skill libraries, agentic tool use, episodic memory and demonstration decomposition are not new, and TGL does not claim them. TGL is the first system to propose the method as a whole: an AI Agent-driven general-purpose robot operating system that connects sparse teaching, explicit closed-loop Skill Blocks carrying a scope and an outcome test, weight-frozen execution, physical feedback and recomposition, structured failure memory, and persistent growth into one single learning cycle. Each neighbouring method covers part of that cycle, so the contribution here is the cycle itself and the interfaces between its parts. That whole-system view is what makes the practical properties available, namely converging faster on a new task, reusing a verified behavior instead of re-deriving it, keeping experience across tasks and embodiments, and serving as a data source for fast policies.

What does TGL stand for?

Teach-and-Grow Learning. The paper is “Teach and Grow: An Agent-Centered Architecture for General Robot Learning”.

How is TGL different from fine-tuning a robot policy?

Fine-tuning changes model parameters to absorb a new behaviour, which can affect previously supported behaviour and requires regression checking. TGL leaves parameters fixed and stores the new capability as an explicit, inspectable Skill Block.

What does a TGL run actually produce?

Two persistent stores: a Skill Library of validated executable behaviours with their scopes and contracts, and an Experience Memory recording the task, the selected blocks, observations, the outcome, the diagnosis and any repair.

How does TGL retain capabilities across tasks?

By writing them into explicit stores instead of into weights. A behaviour that validates on cases kept separate from the teaching demonstrations is admitted to the Skill Library with its scope and outcome test; the conditions, outcome, diagnosis and repair of each attempt go to Experience Memory. Later tasks retrieve from both, so the second attempt at a task starts from what the first one established.

Is a Skill Block just a scripted motion?

No. A scripted motion fixes the trajectory. A Skill Block fixes the intended effect and the conditions, and delegates the motion to a compatible executor, so the same block can run in a different scene with different geometry.

How is a Skill Block different from a function call?

A function call assumes its preconditions hold. A Skill Block states its supported conditions and an outcome test, and the execution loop checks the effect before allowing the next stage.

Can GPT-6 control a robotic arm?

Yes — as a reasoning layer rather than a controller. A GPT-6-class model interprets the visual scene and decides what should happen, then invokes robot-control tools or writes a short program. It does not emit joint commands at control rate; specialist components do that.

What is a GPT-6 robotic arm?

The phrase describes a robot arm driven by a GPT-6-class multimodal model: the model supplies task-level reasoning and tool selection, while perception, grasping and motion come from robot-side components. TGL's implementation uses OpenAI GPT-6 Astra in exactly this role (arXiv:2608.17209).

How do frontier AI models control robot arms?

Three routes appear in the 2026 literature. As a planner, the model produces a sequence that lower layers execute. As a policy, it emits actions directly — the vision-language-action route. As an AI Agent, it stays in the loop, invoking tools or writing programs and revising on physical outcomes. TGL takes the third route.

How does TGL relate to GPT-6 robotic-arm demonstrations?

TGL is one such system, and its report is a worked study of the arrangement: GPT-6 Astra reasons, Codex connects the AI Agent to the robot tools, and the acquisition of new tasks happens outside the weights. The site's paired LIBERO videos show teacher and TGL rollouts on the same tasks.

How is TGL different from direct frontier-model robot control?

Direct control asks the model to produce the action. TGL asks it to produce and check a reusable procedure: each subgoal becomes a Skill Block with an outcome test, and what validates is stored. The model's weights are the same either way; what differs is whether anything persists.

Can TGL work with stronger future multimodal AI Agents?

That is the design intent. Nothing in the architecture depends on this particular model — a stronger AI Agent should ground subgoals better and diagnose failures better, while the Skill Library and Experience Memory carry over unchanged.

Does GPT-6 directly control the arm?

No. It supplies task-level reasoning and tool selection. Joint-level geometry and continuous control come from specialist components, with the model invoking them.

What does TGL add to a GPT-6 robotic arm?

Persistence. TGL wraps each subgoal in a Skill Block with an outcome test, keeps validated blocks in a Skill Library, and records conditions, outcomes, diagnoses and repairs in Experience Memory, so a later task starts from more than a log.

Is Agent as Policy the same as using an LLM for planning?

Not quite. Planning puts the model before execution and commits to a plan. Agent as Policy keeps it running during execution, so it can inspect physical results and revise.

Does Agent as Policy require task-specific training?

No — that is its central claim. Jia et al. demonstrate a general-purpose AI Agent driving a physical robot through task execution with no task-specific or environment-specific training.

Can a general-purpose AI Agent directly control a physical robot?

Yes, and this is the central demonstration of AGP: Jia et al. show a general-purpose AI Agent driving a physical robot through task execution with no task-specific or environment-specific training. What it produces is executable programs and motion commands through a robot interface, not joint torques.

How does TGL relate to Agent as Policy?

TGL adopts the same loop and adds the persistence layer. The AI Agent orders subgoals and revises on physical feedback in both; in TGL, validated behaviour also enters a Skill Library and each attempt's diagnosis enters Experience Memory, so a repeat task starts from what the first one established.

Does the AI Agent produce actions or programs?

Typically programs or tool calls rather than joint targets. In a code-writing variant the AI Agent produces a program that a robot-side layer executes; in a tool-calling variant it selects subgoals and invokes control primitives.

What is a coding AI Agent for robotics?

A model that inspects the robot's state, calls tools, writes and runs a short program, and reads the result. It fits robotics because a robot usually needs a small procedure rather than a single instruction, and because a program can hold intermediate results between steps.

Can Codex-style AI Agents control robots?

They can drive one through tools, but they do not produce control-rate joint commands. In TGL, Codex connects the AI Agent to the robot tools, while detection, RGB-D geometry, Contact-GraspNet, MPLib and controllers do the physical work.

How does TGL relate to coding AI Agents?

Codex is the coding AI Agent in TGL's implementation. What TGL adds is the contract around each call: a subgoal is wrapped in a Skill Block with a stated effect and an outcome test, so running a program is not the same as assuming it worked.

Can a robot learn a task from one video without retraining?

It can acquire the structure of the task that way — the order of operations and the conditions worth checking — and no weights need to change. What the video does not supply is the physical realization: the pose, grasp and motion this scene requires. TGL's answer is to ground each subgoal on the robot and keep what validated.

What is the difference between physical ICL and lifelong robot learning?

Lifelong learning asks how a system keeps acquiring tasks without forgetting; physical ICL asks how a task is acquired without a weight update. They are different axes. TGL sits on both: the acquisition is in-context and the retention is explicit.

Is “without retraining” the same as zero-shot?

No. Zero-shot usually means no task-specific example at all. Learning without retraining still uses a few demonstrations; what it avoids is the policy update, not the teaching.

What replaces the policy update?

An edit to explicit state: a new or narrowed Skill Block, a changed recovery rule, or a record in Experience Memory that changes which block is retrieved next time.

How does robot memory help an AI Agent?

It changes what the AI Agent can do on the second attempt. A diagnosis recorded after one failure lets a later retrieval choose a different grasp family or observation rather than re-running the same plan. Outcome alone — success or failure — is weak evidence; the explanation is the useful part.

What is an experience store for a robot AI Agent?

A record of what was tried and what came of it: the task, the blocks selected, the observations, the outcome, the diagnosis and any repair. In TGL it is Experience Memory, kept separate from the Skill Library so that validated behaviour and contextual history grow independently.