Switch language한국어
Back to the list

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

TL;DR AI

Key summary

2 min read
  1. Researchers released VibeSearchBench, a bilingual benchmark for long-horizon proactive search across 20 domains and 200 curated tasks.

  2. The benchmark uses progressive-disclosure simulations and graph-based evaluation to mirror vague, iterative real-world search behavior.

  3. Testing seven frontier models, including Claude Opus 4.6, GPT-5.4, and Gemini-3.1 Pro, showed low performance and no model reached the user completion signal.

  4. The results expose a major gap between current search-agent benchmarks and real user needs, suggesting proactive search is still unsolved.

Read the original