Statistics Colloquium Series
Date and Time
Location
Our upcoming event for the Statistics Department Colloquium Series is scheduled for Monday, October 16 from 12:00 – 1:00pm (ET) and will be an in-person presentation Science Center Rm. 316. Lunch will be provided to guests following the talk. This week's speaker will be Jacob Steinhardt of the Statistics Department at University of California-Berkeley.
Title: Large Language Models as Statisticians
Abstract: Given their complex behavior, diverse skills, and wide range of deployment scenarios, understanding large language models---and especially their failure modes---is important. Given that new models are released every few months, often with brand new capabilities, how can we achieve understanding that keeps pace with modern practice?
In this talk, I will present an approach to this that leverages the skills of language models themselves, and so scales up as models get better. Specifically, we leverage the skill of language models *as statisticians*. At inference time, language models can read and process significant amounts of information due to their large context windows, and use this to generate useful statistical hypotheses. We will showcase several systems built on this principle, which allow us to audit other models for failures, identify spurious cues in datasets, label the internal representations of models, and factorize corpora into human-interpetable concepts.
This is joint work with many collaborators and students, including Ruiqi Zhong, Erik Jones, and Yossi Gandelsman.
Bio: Jacob is an Assistant Professor of Statistics at UC Berkeley, where he works on trustworthy and human-aligned machine learning. He received his PhD at Stanford University under Percy Liang and has previously worked at OpenAI and Open Philanthropy.