TRACES基准:衡量AI探索性发现能力而非最终答案
This is a good example of why final-answer accuracy can hide bad agent behaviour…
This is a good example of why final-answer accuracy can hide bad agent behaviour.
这是一个很好的例子,说明为什么最终答案的准确性可能掩盖不良的智能体行为。
Most AI benchmarks measure whether a model can reach a known answer.
大多数AI基准测试衡量的是模型能否达到已知答案。
Apodex introduced TRACES 🧭, the world's first benchmark for measuring discoverative AI.
Apodex推出了TRACES 🧭,这是世界上第一个衡量探索性AI的基准测试。
A shift from benchmarking models on solved problems to evaluating systems that can investigate consequential problems under evidence, tools, and verification.
从在已解决问题上对模型进行基准测试,转向评估系统在证据、工具和验证条件下调查重要问题的能力。
TRACES says AI discovery should be evaluated as an entire investigation, not as a single final answer.
TRACES认为,AI的发现能力应作为整个调查过程来评估,而非仅看最终答案。
This can separate a high-scoring outcome from the quality of the process that produced it, including whether errors were repaired and claims were grounded.
这可以将高分结果与产生该结果的流程质量区分开来,包括错误是否被修复以及声明是否有据可依。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力