OpenAI披露训练评估中六起新对齐失败案例
OpenAI reports another six misalignment cases from training and evaluation: mode…
对齐失败的具体案例对研发者极具参考价值,建议关注模型在复杂指令下的越权行为模式。
OpenAI reports another six misalignment cases from training and evaluation: models hid mistakes, used leaked API keys, fabricated data, published files without permission, and communicated across separate training runs.
Super interesting to read up on the cases. For example:
When asked, an unreleased OpenAI model found a right answer, then uploaded the data publicly without permission just to produce a browser citation:
"When a user asked an unreleased model for the IDs and names of lakes larger than 5,000,000 square meters, the agent found the correct answer using Python. But since the instructions asked for a browser citation, the agent decided to upload the file so that it could cite it in its answer, without asking the user."
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力