LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards
TL;DR AI
2 min readKey summary
Researchers introduced LongTraceRL, a reinforcement-learning framework for long-context reasoning in LLMs.
It creates harder training examples from search-agent trajectories and knowledge-graph multi-hop questions, adding tiered distractors.
The method uses entity-level rubric rewards only for correct answers, encouraging more evidence-grounded reasoning.
Across five benchmarks and 4B–30B models, LongTraceRL beat strong baselines on long-context tasks.
