MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
TL;DR AI
2 min readKey summary
Researchers introduced MLS-Bench, a benchmark with 140 tasks across 12 domains to test whether AI systems can invent ML methods that generalize and scale.
The paper finds current AI agents usually lag behind human-designed methods and are better at engineering-style tuning than true method invention.
Results suggest that more compute, search, or longer context alone does not close the gap in scientific discovery ability.
The project also releases data, code, and a community platform to support ongoing evaluation.
