DataSpace基准:强LLM不等于可靠数据Agent
A strong LLM does not automatically make a reliable data agent: DataSpace shows…
A strong LLM does not automatically make a reliable data agent: DataSpace shows that the harness, cross-source joins, and final-table handling all materially affect whether the job actually gets done.
强大的LLM并不自动等同于可靠的数据代理:DataSpace表明,框架、跨源连接以及最终表格处理都会实质性地影响任务是否真正完成。
It tests 410 tasks where agents must combine databases, CSV/JSON files, long documents, and video, then return the exact requested table.
它测试了410项任务,其中代理必须结合数据库、CSV/JSON文件、长文档和视频,然后返回精确请求的表格。
The best model gets only 66.34% right.
最佳模型正确率仅为66.34%。
And with the model fixed, simply changing the agent harness moves accuracy from 30.98% to 46.34%.
而在模型固定的情况下,仅改变代理框架就能将准确率从30.98%提升至46.34%。
Across every tested model, joins and mixing multiple data types are consistent weak spots.
在所有测试的模型中,连接和混合多种数据类型始终是薄弱环节。
But many failures happen even later: the agent has already found or computed the right information, then misunderstands the requested result or submits the wrong columns.
但许多失败甚至发生在更晚阶段:代理已经找到或计算出正确信息,随后误解了请求的结果或提交了错误的列。
– arxiv. org/abs/2608.03451
– arxiv.org/abs/2608.03451
Title: "DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces"
标题:“DataSpace:异构工作空间中可验证分析的数据代理基准测试”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力