Forecasting Downstream Performance of LLMs With Proxy Metrics
TL;DR AI
2 min readKey summary
Researchers propose proxy metrics built from token-level distributions over expert solutions to forecast LLM downstream performance.
The metrics beat cross-entropy loss and compute-based baselines in cross-family model ranking, pretraining data selection, and training-horizon extrapolation.
This could make it cheaper and more reliable to choose architectures, datasets, and training strategies before full downstream evaluation.
