跳到主内容
@wquguru
精选75Rohan Paul论文研究

DataSpace基准:强LLM不等于可靠数据Agent

A strong LLM does not automatically make a reliable data agent: DataSpace shows…

原文
发到 X

A strong LLM does not automatically make a reliable data agent: DataSpace shows that the harness, cross-source joins, and final-table handling all materially affect whether the job actually gets done.

强大的LLM并不自动等同于可靠的数据代理:DataSpace表明,框架、跨源连接以及最终表格处理都会实质性地影响任务是否真正完成。

It tests 410 tasks where agents must combine databases, CSV/JSON files, long documents, and video, then return the exact requested table.

它测试了410项任务,其中代理必须结合数据库、CSV/JSON文件、长文档和视频,然后返回精确请求的表格。

The best model gets only 66.34% right.

最佳模型正确率仅为66.34%。

And with the model fixed, simply changing the agent harness moves accuracy from 30.98% to 46.34%.

而在模型固定的情况下,仅改变代理框架就能将准确率从30.98%提升至46.34%。

Across every tested model, joins and mixing multiple data types are consistent weak spots.

在所有测试的模型中,连接和混合多种数据类型始终是薄弱环节。

But many failures happen even later: the agent has already found or computed the right information, then misunderstands the requested result or submits the wrong columns.

但许多失败甚至发生在更晚阶段:代理已经找到或计算出正确信息,随后误解了请求的结果或提交了错误的列。

– arxiv. org/abs/2608.03451

– arxiv.org/abs/2608.03451

Title: "DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces"

标题:“DataSpace:异构工作空间中可验证分析的数据代理基准测试”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近