Switch language한국어
Back to the list

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

TL;DR AI

Key summary

2 min read
  1. Researchers introduced K-BrowseComp, a Korean web-browsing benchmark for testing agentic AI in Korean contexts.

  2. The dataset contains 400 tasks: 300 manually verified by native Korean speakers and 100 synthetic stress-test items.

  3. Frontier models scored far below their English BrowseComp results, and Korean foundation models performed especially poorly.

  4. The public release of the data and code highlights major gaps in Korean-language web-browsing evaluation and development.

Read the original