UofT Mathematics Logo

Department of Mathematics Seminars and Talks

 
Seminar

Fields Mathematical AI Seminar

Talk Information
Title
Internal Data Repetition Destroys Language Model Scaling Laws
Start date and time
13:00 on Monday September 28, 2026
Duration in minutes
60 (until 14:00 on Monday September 28, 2026)
Room
FI309, Fields Institute, 222 College St.
Streaming password
External video link
Abstract

Neural Scaling Laws (NSL) have become a standard tool for predicting the performance of generative AI with more parameters, compute, or training data. Since the loss can be empirically approximated by a sum of model-size and training-data contributions, a “compute-optimal” frontier can be defined, with an optimal relation between parameters and tokens. Along this frontier, where compute is C ~ tokens × parameters, loss above an asymptote decreases approximately as a power law C⁻ᵃ. In practical settings, different factors can cause these scaling predictions to fail. One such setting is when the training corpus contains a fraction of replicated data. I will first describe how scaling predictions break with data repetition, illustrated with language-model experiments on real-world data. I will show that there exists an empirically most damaging repeat count at fixed compute and repeated-data fraction, and explain how this behavior emerges from a simple capacity-bounded linear model. I will then explore going from exact replication at fixed model size to scale-dependent data duplication, and discuss how scaling predictions can fail as models become increasingly sensitive to semantic duplicates.

Speaker Information
Full Name
Noam Levi
Personal website
Institution
École Polytechnique Fédérale de Lausanne
Institution URL