TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks

TL;DR AI
2 min readKey summary
Researchers introduced TerminalWorld, a benchmark built by reverse-engineering real terminal recordings into authentic command-line tasks.
The team created 1,530 validated tasks and a 200-task verified subset from 80,870 recordings to test AI agents in realistic workflows.
Across models and agents, the best pass rate reached only 62.5%, suggesting terminal performance is still limited.
Scores on TerminalWorld showed weak correlation with existing benchmarks, implying current evaluations may not reflect real-world capability.
