Switch language한국어
Back to the list

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Trajel, a trajectory-level audit framework and dataset for multi-agent industrial workflows.

  2. The study adds a five-part hallucination taxonomy: factual, referential, logical, procedural, and scope-based errors.

  3. Using expert-annotated agent traces from AssetOpsBench, the authors benchmarked detection at multiple levels.

  4. They found that final-answer evaluation misses many failures, and nearly half of hallucinated trajectories contain multiple error types.

  5. Trajectory-aware detectors outperformed standard post-hoc verification, highlighting the need for safer agent deployment.

Read the original