Harvard AI, Math, and Statistics Seminar: Claire Boyer
Date and Time
Location
Our upcoming event for the Harvard AI, Math, and Statistics Seminars is scheduled for Friday, October 9th from 10:30 – 11:30pm (ET) and will be an in-person presentation at Maxwell-Dworkin, Room 134 a/b. This week's speaker will be Claire Boyer, Professor at Paris-Saclay university in the Laboratoire de Mathématiques d'Orsay
How attention learns structure from data?
Transformer architectures have demonstrated remarkable empirical success, yet their ability to extract statistical structure from data remains only partially understood. In this talk, I will present recent work shedding light on attention mechanisms through a statistical and operator-theoretic lens.
Based on large-prompt asymptotics, one can show that softmax attention converges to a linear operator acting on the input-token distribution under Gaussian assumptions. This regime enables a precise analysis of both the outputs and the training dynamics using concentration arguments, revealing a surprising bridge between nonlinear softmax attention and tractable linear models.
These results suggest that attention can be understood as a flexible statistical operator that adapts to the underlying data distribution, providing a unifying framework to study its role in representation learning and in-context inference. I will conclude by discussing the connections between transformer architectures and principal component analysis, using this setting to illustrate the power of concentration methods in high-dimensional statistics.