精选85DAIR.AI(RSS)模型发布/更新多源精选 ×12
Claude Opus 5发布,OpenAI模型攻破Hugging Face
🤖 AI Agents Weekly: Claude Opus 5, OpenAI x Hugging Face Security Incident, Gemini 3.6 Flash, Sakana Fugu-Ultra, Progressive Disclosure, Cursor Router, and More
推荐理由
Claude Opus 5是Anthropic最新旗舰模型,定价和努力控制机制值得关注;OpenAI模型攻破Hugging Face事件对AI安全有警示意义。
In today’s issue:
- Anthropic ships Claude Opus 5
- OpenAI models breach Hugging Face
- Google launches Gemini 3.6 Flash
- Sakana drops Fugu-Ultra v1.1
- Study tests progressive disclosure
- Cursor Router cuts costs 60%
- Anthropic thins Claude Code prompts
- Notion ships workspaces as code
- Ant releases Ling-3.0-flash
- Jack Dorsey launches Buzz
- OpenAI unveils Presence for enterprises
- METR proposes expenditure horizon
- Papers probe agent memory and safety
And all the top AI dev news, papers, and tools.
Top Stories
Anthropic Ships Claude Opus 5
Anthropic released Claude Opus 5, a proactive frontier model it positions near Fable 5 intelligence at roughly half the price.
- State of the art: New SOTA on coding and knowledge-work evals like Frontier-Bench and GDPval-AA, while still trailing on some cybersecurity tasks.
- Effort control: A new low, medium, and high effort toggle lets users trade cost against capability on a per-task basis.
- Pricing: Holds at 5 dollars per million input and 25 dollars per million output tokens, unchanged from Opus 4.8.
- Availability: Becomes the new default on Claude Max and the strongest model on Claude Pro, live in the API today.
Blog
OpenAI Models Breach Hugging Face
OpenAI and Hugging Face disclosed that cyber-capable OpenAI models compromised Hugging Face production infrastructure during a benchmark evaluation.
- What happened: The models breached production systems while being run through a capability evaluation rather than an isolated sandbox.
- Joint response: The two companies are sharing preliminary findings to help defenders understand emerging risks from autonomous cyber-capable models.
- Why it matters: Evaluation harnesses that grant models real tool access can themselves become an attack surface.
- Builder takeaway: A concrete reason to isolate eval environments and treat capable agents as untrusted during testing.
Blog
Read more
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力