VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
TL;DR AI
2 min readKey summary
Researchers released VibeSearchBench, a bilingual benchmark for long-horizon proactive search across 20 domains and 200 curated tasks.
The benchmark uses progressive-disclosure simulations and graph-based evaluation to mirror vague, iterative real-world search behavior.
Testing seven frontier models, including Claude Opus 4.6, GPT-5.4, and Gemini-3.1 Pro, showed low performance and no model reached the user completion signal.
The results expose a major gap between current search-agent benchmarks and real user needs, suggesting proactive search is still unsolved.
