精选85The Decoder(RSS)行业动态多源精选 ×6
OpenAI用AI攻击自家AI,成功率84%远超人类
OpenAI is now using AI to attack its own AI, and it's working better than humans ever did
推荐理由
做AI安全和对齐的同学重点关注,GPT-Red的自我对抗训练方法可能成为行业新范式,建议研究其技术细节并评估自身模型的安全测试流程。
OpenAI's internal GPT-Red model finds successful attacks in 84 percent of test scenarios through self-play training. Human red teamers manage just 13 percent. The results feed directly into hardening models like GPT-5.6 Sol. The article OpenAI is now using AI to attack its own AI, and it's working better than humans ever did appeared first on The Decoder.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力