Synthetic Sandbox for Training Machine Learning Engineering Agents
TL;DR AI
2 min readKey summary
SandMLE is a multi-agent framework that creates synthetic MLE environments from a few seed tasks to enable efficient on-policy RL.
By constraining datasets to 50–200 samples per task, SandMLE reduces execution time by over 13× compared with full-scale pipelines.
Using SandMLE, experiments show larger gains than supervised fine-tuning on MLE-bench-lite across Qwen3 series models.
The trained policy also generalized to new agent scaffolds, improving HumanRank on MLE-Dojo by up to 32.4%.
