跳到主内容
@wquguru
精选70OpenAI产品发布/更新多源精选 ×2

OpenAI发布模型不对齐行为追踪与披露框架

We're sharing our new framework for tracking, investigating, and disclosing inst…

原文
发到 X

We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.

我们正在分享我们在 OpenAI 用于追踪、调查和披露模型不对齐实例的新框架。

The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties.

该框架为公开披露设定了标准和时间表,包括在我们尚未完全解释或缓解该行为时的情况。更复杂的案例可能需要更长的调查时间或与第三方协调。

We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation.

我们将优先考虑那些揭示新不对齐机制、已知行为的重大变化,或对安全或缓解措施假设提出挑战的发现示例。

Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months.

与该框架一同发布的,还有六份关于我们在过去六个月中在训练或评估模型期间观察到的不对齐行为实例的报告。

This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.

这是一个起点。我们将通过经验和公众反馈来完善这一流程,并持续分享更多报告。

https://openai.com/index/model-misalignment-reporting-framework/

https://openai.com/index/model-misalignment-reporting-framework/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →