ResearchMath-14K: Scaling Research-Level Mathematics via Agents
TL;DR AI
2 min readKey summary
Researchers introduced ResearchMath-14K, a 14,056-problem dataset of research-level math built from academic sources using a multi-agent pipeline.
They also created ResearchMath-Reasoning, 220K reasoning trajectories from two open models, to study how models approach hard math problems.
The team found frequent failures such as not attempting a solution or inventing citations, then filtered the trajectories to remove low-quality examples.
Using the filtered data to fine-tune Qwen3 models from 4B to 30B parameters improved performance by 9.2 points on average.
The work offers the largest public collection of research-level math problems and shows that imperfect attempts can still help train stronger reasoning models.
