Agent as Policy (AGP)

Agent as Policy (AGP) places a general-purpose AI Agent inside the execution loop rather than limiting it to offline planning. The AI Agent observes the robot and environment, reasons at runtime, invokes control tools or executable programs, inspects the physical result, and revises its next action. Jia et al. named the approach and demonstrated it on real manipulation tasks in September 2026.

Short answer

Agent-as-Policy robotics places a general-purpose AI Agent inside the execution loop rather than limiting it to offline planning. The AI Agent observes the robot and environment, reasons at runtime, invokes control tools or executable programs, inspects the physical result, and revises its next action.

Where the name comes from

“Agent as Policy for Robotic Manipulation” (Jia et al., arXiv:2609.12541, September 2026) introduces AGP and shows a general-purpose AI Agent driving a physical robot through task execution with no task-specific or environment-specific training. Given a task and a robot interface, the AI Agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes — across precision manipulation, dynamic motions and deformable-object tasks.

The contrast is with the foundation-policy line, where a vision-language-action model maps observations to actions directly. AGP keeps the AI Agent running in the loop and gives it programs and tools rather than joint targets.

Key idea

What distinguishes the placement is where the deciding component sits: outside the loop, producing a plan that lower layers execute, or inside it, reacting to what the robot actually observed. Physical execution produces evidence — a failed grasp, an object that moved, a drawer that stayed shut — and only a component that is still running can act on it.

The cost is latency and cost per step. Reasoning on every action is far more expensive than feed-forward inference, which is why the placement tends to be reserved for novelty: unfamiliar objects, diagnosis and recovery.

Current examples

Agent as Policy (AGP) — Jia et al., arXiv:2609.12541, 2026: a general AI Agent drives a real robot across manipulation tasks with no task-specific training.

Agentic Robot — Yang et al., arXiv:2505.23450, 2025: a framework for vision-language-action models that adds an action-coordination protocol and execution-time verification for long-horizon manipulation.

Push-T with agentic robotics — Xie, Chen and Goldberg, arXiv:2608.18227, 2026: an LLM coding AI Agent writes a solution to Push-T with no demonstration data, compared against a visuomotor imitation policy.

Code as Policies — Liang et al., arXiv:2209.07753, 2022: the program-as-policy predecessor, where a language model writes policy code over perception primitives.

SayCan — Ahn et al., arXiv:2204.01691, 2022: grounding language-model plans in what the robot can actually do.

ReKep — Huang et al., arXiv:2409.01652, 2024: relational keypoint constraints for closed-loop manipulation — a spatial-reasoning route to the same problem.

Agent as Policy compared with Teach-and-Grow Learning

The two share the control locus — an AI Agent inside the loop — and differ on what persists. Across the dimensions that matter for a robot acquiring tasks over time:

Control locus: identical. Both keep the AI Agent inside the execution loop.

Task acquisition: AGP acquires from the task description and the robot interface; TGL acquires from sparse demonstrations, which supply subgoal structure and the conditions worth checking.

Runtime reasoning: identical. Both reason while the task runs.

Skill persistence: AGP does not define a persistent store; in TGL, validated behaviour enters a Skill Library as Skill Blocks.

Memory: AGP carries state within the task; TGL keeps a separate Experience Memory of outcome, diagnosis and repair across tasks.

Experience reuse: AGP re-derives a solution on a repeat task; TGL retrieves the validated block.

Demonstration use: AGP requires none; TGL uses a few.

Task-specific retraining: neither updates the policy — this is the shared claim.

Physical feedback: both inspect the physical outcome; TGL makes the effect check part of each Skill Block's contract.

Future-task transfer: TGL's explicit claim and the report's scaling hypothesis; not a claim AGP makes.

How TGL relates

TGL follows the Agent-as-Policy design and adds the persistence layer. The AI Agent orders subgoals, chooses tools and revises the route; validated behaviour accumulates in a Skill Library, and the conditions, outcomes, diagnoses and repairs of each attempt accumulate in Experience Memory. The report's slow-teacher/fast-student split lets a learned policy take over mature behaviours, keeping agentic deliberation for novelty.

Frequently asked questions

Is Agent as Policy the same as using an LLM for planning?

Not quite. Planning puts the model before execution and commits to a plan. Agent as Policy keeps it running during execution, so it can inspect physical results and revise.

Does Agent as Policy require task-specific training?

No — that is its central claim. Jia et al. demonstrate a general-purpose AI Agent driving a physical robot through task execution with no task-specific or environment-specific training.

Can a general-purpose AI Agent directly control a physical robot?

Yes, and this is the central demonstration of AGP: Jia et al. show a general-purpose AI Agent driving a physical robot through task execution with no task-specific or environment-specific training. What it produces is executable programs and motion commands through a robot interface, not joint torques.

How does TGL relate to Agent as Policy?

TGL adopts the same loop and adds the persistence layer. The AI Agent orders subgoals and revises on physical feedback in both; in TGL, validated behaviour also enters a Skill Library and each attempt's diagnosis enters Experience Memory, so a repeat task starts from what the first one established.

Does the AI Agent produce actions or programs?

Typically programs or tool calls rather than joint targets. In a code-writing variant the AI Agent produces a program that a robot-side layer executes; in a tool-calling variant it selects subgoals and invokes control primitives.

Related pages