Physical AI and Embodied AI
Physical AI and embodied AI are the terms now used for systems that perceive and act in the physical world through sensors and actuators, rather than producing text or images. Behind the labels is a set of concrete engineering problems: how a policy's sensory interface should be shaped, how it behaves when decisions are delayed, and where its training data comes from.
The label and the substance
Physical AI is the industry framing; embodied AI and embodied intelligence are the more academic ones. All three point at the same shift. A language model predicts the next token; a physical AI Agent has to deal with the consequences of its own actions in a world that pushes back. Mass, friction, inertia and contact are not in the training distribution of text, so the representation a physical AI Agent needs is not the one a chatbot needs.
What is genuinely hard
Data. Internet text and video are abundant and third-person. A robot needs first-person evidence of what the world becomes after it acts, and that data is expensive to create.
Sensory interfaces. Most work assumes vision suffices. It does not: contact, force and hidden internal state are invisible to cameras, and the modalities that do report them each carry their own temporal structure.
Timing. Physical AI Agents run under latency and their most capable policies run slowly. Anything that must be noticed between two decisions falls into the gap.
Evaluation. Reaching a goal is not the same as behaving correctly. A policy can look successful on a geometric metric while being wrong in the way that matters.
Where TGL fits
Teach and Grow is a physical-AI system. Its subject is the AI Agent's interface to the physical world: a robot acquiring a new manipulation capability by acting, observing the outcome, and keeping what validated. Its claim is that for AI Agents under delayed control, preserving what happened is a requirement rather than a refinement.