UofT Mathematics Logo

Department of Mathematics Seminars and Talks

 
Seminar

Fields Mathematical AI Seminar

Talk Information
Title
Corrigibility Transformation: Constructing Goals That Accept Updates
Start date and time
13:00 on Monday September 21, 2026
Duration in minutes
60 (until 14:00 on Monday September 21, 2026)
Room
FI309, Fields Institute, 222 College St.
Streaming password
External video link
Abstract

An AI agent will learn a desired goal more effectively if it does not resist the training process, but many partially learned goals incentivize an AI to avoid further goal updates. We would like goals to be corrigible, meaning they allow changes requested through designated channels, so that we can confidently correct errors and shut down the AI if necessary. Despite this being a crucial safety property, the existing literature does not specify goals that are both corrigible and competitive with alternatives. We introduce a transformation that constructs a corrigible version of nearly any goal, without sacrificing performance. This is done by eliciting predictions of reward conditional on costlessly preventing updates, and having that target be pursued myopically. These goals are then shown to lead to optimal performance among the class of corrigible goals, to incentivize allowing mid-action overrides, and to disincentivize deliberate self-modification. Empirically, they induce corrigible behavior in gridworld settings and for language models when applied at the prompt level.

Speaker Information
Full Name
Rubi Hudson
Personal website
Institution
University of Toronto
Institution URL