Fields Mathematical AI Seminar
by Mikhail Belkin (University of California, San Diego)
A trained Large Language Model (LLM) contains much of human knowledge. Yet, it is difficult to gauge the extent or accuracy of that knowledge, as LLMs do not always ``know what they know'' and may even be unintentionally or actively misleading. In this talk I will discuss feature learning and some interesting behaviors of MLPs, seemingly related to LLMs. I will introduce Recursive Feature Machines—a powerful method originally designed for extracting relevant features from tabular data. I will show how this technique enables us to detect and guide LLM behaviors toward almost any desired concept by adding a multiple of fixed vectors in LLM activation spaces. Finally I will discuss a few, perhaps only tangentially related, thoughts on alignment.