K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
TL;DR AI
2 min readKey summary
Researchers introduced K-BrowseComp, a Korean web-browsing benchmark for testing agentic AI in Korean contexts.
The dataset contains 400 tasks: 300 manually verified by native Korean speakers and 100 synthetic stress-test items.
Frontier models scored far below their English BrowseComp results, and Korean foundation models performed especially poorly.
The public release of the data and code highlights major gaps in Korean-language web-browsing evaluation and development.
