HLL: Can Agents Cross Humanity's Last Line of Verification?

TL;DR AI
2 min readKey summary
Researchers introduced HLL, a GUI benchmark built around interactive CAPTCHA tasks to test whether multimodal agents can cross human-verification boundaries.
Across eight frontier agents, performance was brittle and varied sharply by CAPTCHA type, with harder or cluttered interfaces causing bigger drops.
Agents did even worse when solutions had to be supported by valid action traces, underscoring limits in closed-loop, security-sensitive workflows.
