The Retraining Tax

The retraining tax is the recurring cost of repairing robot behaviour through a policy update. It includes new data collection, optimisation, and regression checking against everything the policy already supported. The term is introduced in the Teach and Grow paper, which uses it to frame an alternative: store the new capability explicitly and leave the weights alone.

Why the cost recurs

End-to-end policies absorb a new behaviour into shared parameters. That is what makes them general, and it is also why a local failure rarely has a local fix: adding corrective data and changing the parameters does not produce a separately addressable repair for one object or one contact condition. Previously supported behaviour may need to be checked again. The same applies when a sensor is added or a gripper changed, which introduces observation interfaces, calibration and action compatibility to validate.

Why the long tail exposes it

The cost is tolerable while changes are broad and infrequent. It becomes visible in the long tail, where a rare contact condition or an unusual object needs a specific lesson rather than another wide round of experience. The cheaper an individual correction should be, the more the shared-parameter route costs relative to it.

What TGL does instead

TGL keeps the pretrained stack fixed and stores new capability in explicit objects: Skill Blocks with stated scopes and outcome tests, and Experience Memory of the conditions and repairs. A correction becomes an edit to one of those objects. The paper analyses when this is genuinely cheaper — grounding, validation, compatibility checking and retrieval all have to stay manageable — rather than assuming it always is.

How the term should be used

The retraining tax is a framing device for a cost structure, not a measured quantity in the paper. It is useful for asking, of any robot-learning system, what has to be redone when one behaviour is repaired. It should not be quoted as an empirical measurement.

Related pages