跳到主内容
@wquguru
精选70Rohan Paul论文研究

AI Agent 在真实 CAPTCHA 验证中仍表现不佳

Today’s AI agents still struggle to pass real human-verification checks (CAPTCHA…

原文
发到 X

Today’s AI agents still struggle to pass real human-verification checks (CAPTCHAs) on websites.

The paper proposes HLL, a benchmark where agents must solve 10 types of CAPTCHA tasks by seeing the page, clicking or dragging correctly, tracking state, and submitting the answer.

A useful agent must find the right box on a messy page, understand the instruction, click or drag in the right place, track what changed, recover from mistakes, and leave an interaction trail that looks consistent with the task.

The paper shows that even strong agents can look smart on static tasks, then fail when the page is cluttered, the task is harder, or the system checks whether their actions were actually valid.

Link – arxiv. org/abs/2606.02449

Title: "HLL: Can Agents Cross Humanity's Last Line of Verification?"

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近