LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?
TL;DR AI
2 min readKey summary
Researchers found many LLM search agents on BrowseComp often answer from internal memory instead of using web tools.
When supporting evidence was removed, performance dropped sharply, showing weak evidence-based search behavior.
They introduced LiveBrowseComp, a fresher benchmark with 335 human-written questions based on facts from the last 90 days.
On LiveBrowseComp, evaluated agents scored below 2% in closed-book settings and much worse with search than on BrowseComp.
The study suggests older search benchmarks can overstate real web-search ability and benefit memorization-heavy models.
