UofT Mathematics Logo

Department of Mathematics Seminars and Talks

 
Seminar

Fields Mathematical AI Seminar

Talk Information
Title
Feature learning, alignment and the linear representation hypothesis for steering and monitoring LLMs
Start date and time
14:00 on Wednesday September 02, 2026
Duration in minutes
60 (until 15:00 on Wednesday September 02, 2026)
Room
FI230, Fields Institute, 222 College St.
Streaming password
External video link
Abstract

A trained Large Language Model (LLM) contains much of human knowledge. Yet, it is difficult to gauge the extent or accuracy of that knowledge, as LLMs do not always ``know what they know'' and may even be unintentionally or actively misleading. In this talk I will discuss feature learning and some interesting behaviors of MLPs, seemingly related to LLMs. I will introduce Recursive Feature Machines—a powerful method originally designed for extracting relevant features from tabular data. I will show how this technique enables us to detect and guide LLM behaviors toward almost any desired concept by adding a multiple of fixed vectors in LLM activation spaces. Finally I will discuss a few, perhaps only tangentially related, thoughts on alignment.

Speaker Information
Full Name
Mikhail Belkin
Institution
University of California, San Diego
Institution URL