LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis
TL;DR AI
2 min readKey summary
LongDS benchmarks AI agents on long-horizon data analysis using 68 real Kaggle notebook tasks across six domains and 2,225 turns.
Across five leading models, the best result was 48.45% accuracy, showing limited reliability in extended analytical workflows.
Performance fell sharply as conversations progressed, with most mistakes caused by failures to track and revise analytical state.
The findings highlight a major gap in current agents’ ability to support dependable, realistic data science work.
