Omar Sar:当前模型评测基准存在缺陷,Harness工程成关键
Important discussion. Measuring models against harnesses is completely broken. I…
Important discussion. Measuring models against harnesses is completely broken. I prefer to test model quality against minimal harnesses like Pi and Hermes Agent. This is not perfect, as there are biases in the harnesses that favor some models and not others.
重要的讨论。用基准测试来衡量模型是完全错误的。我更喜欢用像 Pi 和 Hermes Agent 这样的最小化基准来测试模型质量。这并不完美,因为基准中存在偏向某些模型而非其他模型的偏差。
A standard way to do this is missing but important, as harness engineering is where leading AI companies are focusing efforts. Not enough effort here, as things are moving fast and harness optimization is individualistic.
目前缺少一种标准的做法,但这很重要,因为基准工程是领先 AI 公司正在投入精力的领域。这里的努力还不够,因为技术发展迅速,且基准优化具有个人主义色彩。
On the flip side, I feel like models will eventually have the ability to dynamically generate harnesses on the fly as per task. Claude models do this already to some extent, though pretty inconsistently and remain a mystery. But this could mean that a harness is just a tunable artifact like a system prompt. In that realm, how are we assessing it, and exactly what? Benchmarking will only get murkier from here onwards.
另一方面,我觉得模型最终将具备根据任务动态生成基准的能力。Claude 模型已经在某种程度上做到了这一点,尽管相当不一致且仍是个谜。但这可能意味着基准只是一个可调的产物,就像系统提示词一样。在这个领域,我们如何评估它,以及具体评估什么?从今往后,基准测试只会变得更加模糊不清。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力