Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows
TL;DR AI
2 min readKey summary
Researchers introduced Trajel, a trajectory-level audit framework and dataset for multi-agent industrial workflows.
The study adds a five-part hallucination taxonomy: factual, referential, logical, procedural, and scope-based errors.
Using expert-annotated agent traces from AssetOpsBench, the authors benchmarked detection at multiple levels.
They found that final-answer evaluation misses many failures, and nearly half of hallucinated trajectories contain multiple error types.
Trajectory-aware detectors outperformed standard post-hoc verification, highlighting the need for safer agent deployment.
