Switch language한국어
Back to the list

HLL: Can Agents Cross Humanity's Last Line of Verification?

TL;DR AI

Key summary

2 min read
  1. Researchers introduced HLL, a GUI benchmark built around interactive CAPTCHA tasks to test whether multimodal agents can cross human-verification boundaries.

  2. Across eight frontier agents, performance was brittle and varied sharply by CAPTCHA type, with harder or cluttered interfaces causing bigger drops.

  3. Agents did even worse when solutions had to be supported by valid action traces, underscoring limits in closed-loop, security-sensitive workflows.

Read the original