跳到主内容
@wquguru
精选70elvis论文研究

新基准DataSpace:数据Agent跨语言任务准确率仅66%

Harness choice is a big deal.

原文
发到 X

Harness choice is a big deal.

So much room to advance and improve results across the board with agent harnesses.

Great paper highlighting this.

New research releases DataSpace, a benchmark where data agents produce verifiable tabular results from heterogeneous workspaces. 410 cross-language tasks over 7,439 artifacts totalling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video.

Across six recent frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%. With the backbone held fixed, swapping the harness moves accuracy by 15.36 points.

Multimodal evidence integration and joins reduce accuracy across all six backbones. The benchmark is nowhere near saturated.

Paper: https://arxiv.org/abs/2608.03451

Track more trending AI papers in our academy: https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近