Tool Use in Robotics

Tool use in robotics means exposing perception, grasping, motion and control as callable capabilities that a reasoning AI Agent selects and invokes. It is the mechanism that lets a model act on the physical world without emitting control signals itself.

Why tools rather than one model

A robot needs capabilities that a language model does not have: metric depth, collision-free motion, contact, and control at tens of hertz. Wrapping each as a tool keeps the AI Agent's job at the level it is good at — deciding what should happen — and keeps the geometry where it belongs.

It also makes the system inspectable. A tool has a documented effect and a version; when a behaviour stops working, the question of which component changed has an answer.

What a good tool reports

The return value matters more than the call. A tool that reports “command sent” gives the AI Agent nothing to reason with. A tool that reports the intended physical effect and whether it was observed gives the AI Agent something it can act on — and gives the whole system a place to check causality rather than assume it.

This is the same argument as the outcome test on a Skill Block, seen from the tool side.

Where tools and skills meet

A skill is what the robot can do; a tool is how it does it. TGL's Skill Block is explicit about the relationship: a block declares which executors can realize it and what evidence counts as success. That declaration is what makes a block portable across tool versions rather than bound to one implementation.

How TGL relates

The AI Agent selects subgoals and invokes tools; detection, segmentation, RGB-depth geometry, Contact-GraspNet, MPLib and controllers supply the physical operations, with Codex connecting the AI Agent to them. What TGL adds is the contract around each call — a stated effect and a test — so that invoking a tool is not the same as assuming it worked.

Related pages