Fields Mathematical AI Seminar
by Noam Levi (École Polytechnique Fédérale de Lausanne)
Neural Scaling Laws (NSL) have become a standard tool for predicting the performance of generative AI with more parameters, compute, or training data. Since the loss can be empirically approximated by a sum of model-size and training-data contributions, a “compute-optimal” frontier can be defined, with an optimal relation between parameters and tokens. Along this frontier, where compute is C ~ tokens × parameters, loss above an asymptote decreases approximately as a power law C⁻ᵃ. In practical settings, different factors can cause these scaling predictions to fail. One such setting is when the training corpus contains a fraction of replicated data. I will first describe how scaling predictions break with data repetition, illustrated with language-model experiments on real-world data. I will show that there exists an empirically most damaging repeat count at fixed compute and repeated-data fraction, and explain how this behavior emerges from a simple capacity-bounded linear model. I will then explore going from exact replication at fixed model size to scale-dependent data duplication, and discuss how scaling predictions can fail as models become increasingly sensitive to semantic duplicates.