LLM Robotics

LLM robotics is the use of large language models in a robot's control stack. What they contribute is task decomposition, tool selection, state tracking and recovery — reasoning over symbols and outcomes. What they cannot contribute is a control signal, because they do not run at anything close to the frequency a robot needs.

What an LLM is actually good for here

Decomposing “make coffee” into ordered steps is largely a symbolic problem, and a language model handles it well. So does choosing which tool to invoke next, noticing that an outcome contradicts the plan, and proposing an alternative. These are the parts of a task that benefit from a long context and slow deliberation.

Multimodal models extend this to reading the scene. An AI Agent that can look at a camera frame and see that the drawer did not open has something concrete to reason about, rather than reasoning only over a text description of the state.

What an LLM cannot do

It cannot produce joint torques, and re-running a large model before every low-level command is not feasible at control frequency. In practice the two layers are separated: the AI Agent decides what and the robot-side components handle how.

This also means the AI Agent only knows what its tools tell it. A low-level interface that discards information constrains the AI Agent's reasoning, not just its control.

How TGL uses one

In Teach-and-Grow Learning the multimodal AI Agent identifies subgoals shared across demonstrations, expresses them as closed-loop Skill Blocks, and revises the remaining plan from what the robot observed. The paper's implementation uses OpenAI GPT-6 Astra for that reasoning, with Codex connecting the AI Agent to the robot tools. Detection, segmentation, RGB-D geometry, Contact-GraspNet, MPLib and controllers supply the physical grounding.

Related pages