METR introduces a new metric to calculate exactly when AI agents become more expensive than humans

TL;DR AI
2 min readKey summary
METR introduced a new “expenditure horizon” metric to compare AI agents and human labor on equal performance gains.
Applied to the NanoGPT speedrun, the study estimated human effort at about $2,500 per 1% speedup.
Several AI models delivered only modest gains, with expenditure horizons ranging from $0 to $3,300.
Older models often failed to outperform noise, while newer ones varied widely in cost-effectiveness.



