Agentic Robotics
Agentic robotics means putting a reasoning AI Agent — typically a large multimodal model — in charge of what a robot should do next, while specialised components handle geometry and continuous control. The AI Agent reads the scene, chooses a subgoal and a tool, observes the result, and revises the plan. Teach-and-Grow Learning (TGL) is an AI Agent-centered architecture built on exactly that loop.
The loop
An agentic system is defined less by its model than by its cycle: observe, decide, act, read the outcome, revise. In a workspace that means camera observations standing in for state, perception and motion tools performing physical operations, and executor reports describing what happened. The AI Agent's next choice depends on those reports rather than on a fixed script.
This is what made tool-using language AI Agents useful in software, applied somewhere less forgiving: a mistaken action changes the physical world, and the correction has to come from what the robot actually observed.
Why the division of labour matters
A language model cannot emit joint torques, and a manipulation policy cannot reason about a multi-minute task. The two have incompatible requirements — long context and slow deliberation on one side, tens of hertz on the other. Agentic robotics splits them rather than trying to make one model do both, which produces something either too slow to control or too shallow to plan.
Where the AI Agent stops being able to help
An AI Agent reasons over the world it can see, and it sees only what its tools report. If the low-level interface silently discards something — a brief event that fell between two decisions, say — then that information is missing from the AI Agent's model of the world too, and its plan is built on an incomplete picture. This is why the low-level interface is an AI Agent-level concern rather than only a control detail.
Where TGL fits
TGL is AI Agent-centered by design. The AI Agent identifies subgoals shared across demonstrations, expresses them as closed-loop Skill Blocks, grounds each block in the current scene, and decides what to keep. The robot-side executors supply the geometry and control. In the paper's implementation the reasoning comes from OpenAI GPT-6 Astra, with Codex connecting the AI Agent to the robot tools.