Colloquium Series: Eugene Katsevich
Date and Time
Location
Our upcoming event for the Statistics Colloquium Series is scheduled for Monday, February 23 from 12:00 – 1:00pm (ET) and will be an in-person presentation at Maxwell-Dworkin 134A/B. Lunch will be provided to guests following the talk. This week's speaker will be Eugene Katsevich of UPenn Wharton's Department of Statistics and Data Science.
Power of masking methods for adaptive testing in a multivariate normal means problem
Many large-scale testing procedures learn signal structure from the data to boost power. Direct data reuse can inflate Type-I error ("double dipping"), so a common remedy is masking: withholding some information during learning and using it for testing. Sample splitting masks by withholding observations for testing, while null augmentation (e.g., knockoffs or full-conformal outlier detection) masks by appending null samples or variables and withholding their identities until testing. Little is known about how the power of masking methods compares across mechanisms, across tuning choices, or against more data-efficient non-masking alternatives. Consequently, mechanism selection and tuning are often guided by heuristics with potentially substantial power consequences. Working within a two-groups multivariate normal means model with an unknown signal direction learned from the data, we develop a transparent, unified set of asymptotic power expressions for Split BH, BONuS (closely related to full-conformal outlier detection), and an oracle in-sample benchmark that learns and tests on the full data. Our main findings are: (1) BONuS is more powerful than Split BH across tuning choices; (2) the power-optimal number of null samples for BONuS is a vanishing fraction of the number of tests, in which case its power approaches that of the in-sample benchmark; and (3) for a tractable approximation to BONuS, the optimal number of null samples scales as the square root of the number of tests, with empirical evidence suggesting a similar scaling for BONuS itself. These results characterize masking-induced power tradeoffs in a tractable model and suggest qualitative lessons for mechanism choice and tuning more broadly.
Eugene Katsevich is an assistant professor at Wharton’s Department of Statistics and Data Science. His research centers on several interconnected themes, spanning statistical theory, methodology, and applications. He works on modern genomics applications, which inspire him to consider methodological and theoretical problems in high-dimensional variable selection, conditional independence testing, double/debiased machine learning, multiple testing, and selective inference.