The Distillation Game: Adaptive Attacks & Efficient Defenses

TL;DR AI
2 min readKey summary
A machine-learning paper reframes model distillation as a minimax game between a utility-constrained teacher and an adaptive student.
It introduces adaptive evaluation methods and a forward-pass-only Product-of-Experts defense against model extraction and imitation.
The study finds adaptive students recover much more capability than passive tests suggest, weakening confidence in standard evaluations.
On GSM8K and MATH, the cheaper Product-of-Experts defense performs nearly as well as more expensive defenses under stronger testing.
