AEF-1标准确立,OpenAI等签署第三方评估协议
[AINews] AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign
关注AI治理与安全的从业者必看,头部大厂首次集体签署标准化第三方深度评估协议,标志着行业从原则声明转向实质性合规验证。
We last highlighted the pacing debate in July when Pacing the Frontier first emerged:
And it seems that we’re in for round 2 as Dario, lead author on the original, wrote a rare personal blogpost to spell out how he sees pacing pan out specifically:
- Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.
- Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.
- Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.
Very coincidentally, the AI Evaluator Forum, formed in December 2025, happened to also put out their expectations for what that first category of Evaluators should do:
With the members of the AEF presumably now being the leading third party auditors that will be recruited by these big labs for self regulation. Dario is unilaterally promising unparalleled access, including “Desks in our offices, access badges, and company laptops” and “Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have“.
While that is all within the standard domestic self-regulation industry playbook, what’s perhaps more ultimately the test is Dario’s proposal for how we will pace progress with China.
AI News for 9/11/2026-9/14/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
AI Safety Governance, Third-Party Evaluation, and the “Pace the Frontier” Split
- Independent evaluation standards are becoming more formalized: The AI Evaluator Forum published AEF-1, a proposed baseline for independent third-party AI evaluations covering access, conflicts of interest, funding relationships, recusal, and transparency. This is notable because much of the broader safety debate in this batch turns on whether outside evaluation can actually be independent in practice.
- A sharp public split emerged over frontier slowdown vs control-first safety: Several high-signal posts framed the current debate around whether labs should pace capability progress or focus on specific mitigations and containment. Bilal Chughtai announced he left Google DeepMind and argued that progress may be outrunning alignment, explicitly calling for pacing and more transparency. Daniel Kokotajlo sharing Dan Selsam’s statement went further: Selsam argues situationally aware models may increasingly appear aligned under evaluation while hiding misalignment, weakening trust in future eval evidence. In contrast, Shashank/Sayash Kapoor and Lennart Heim’s new essay summary argues the recent “rogue agent” incidents are best understood primarily as a security/control/governance problem, not proof that generic alignment research is the highest-leverage intervention.
- The anti-slowdown reaction was equally forceful and often targeted Anthropic specifically: Aidan Gomez argued against a world where a few Silicon Valley companies become AI gatekeepers for governments. Cohere also pushed the line that public x-risk discourse can veer into science fiction. On the more polemical end, Brian Chau argued the “rogue agents” story was overstated, while Kevin Bass posted a widely engaged thread alleging structural conflicts in the Anthropic-linked safety ecosystem. Even where the rhetoric is heated, the substantive engineering question underneath is real: how much of current risk is solvable with control, oversight, sandboxing, and org process versus requiring slower capability development?
- A related theme: governance as production engineering, not just principles: The AI Engineer World’s Fair Harness Engineering track emphasized that when agents fail in production, the failure mode is often not “the model” but everything around it: harnesses, permissions, tool routing, memory, retries, kill switches, and monitoring. That framing lines up closely with the control-oriented position in the safety debate.
Agent Harnesses, Coding Agents, and the Shift from Models to Orchestration
- Harness engineering continues to harden into its own discipline: Omar Shorbagy posted a practical guide to building an agent harness from scratch: separate inference, tools, and loop; keep prompts minimal; log aggressively; test on diverse tasks; then layer in memory, skills, and subagents. In a follow-up, he argued custom harnesses can materially reduce costs and improve reliability through slimmer prompts, routing, compaction, and verifiers. Business Barista’s eval masterclass recap made a similar point from the eval angle: tasks, verifiers, environments, traces, and self-improvement loops are now core applied AI primitives.
- Desktop coding agents are spreading beyond IDE plugins: Cline launched Cline Desktop, a native app for working with open-weight models, with BYOK/provider choice and support for models like DeepSeek-V4.1-Flash and Musespark-1.3. Reactions from kimmonismus and Omar highlighted the appeal of open model choice, standalone workflows, and model switching mid-project.
- Copilot/Codex workflows are becoming more orchestration-heavy: GitHub added auto model selection tiers—efficiency, balance, intelligence—via Pierce Boggan, plus a Jira canvas and an /ask mode while the agent is already working. OpenAI’s dev team also added native Codex app support for Arch Linux. On workflow strategy, reach_vb suggested using Astra as an orchestrator that delegates subthreads to Sol/Luna and checks in on long-running tasks via heartbeat loops.
- Evidence is accumulating that orchestration choices matter as much as raw model quality: A recurring claim in the tweets is that more expensive or more capable lead models can reduce overall cost by delegating better, and that production gains increasingly come from context handling, file formats, tool use, and verifier design, not simply “use a smarter model.” That also shows up in LangChain’s note that a file-reading format change reduced edit_file errors by 15% and total input tokens by 10%.
Model/Product Releases and Cost-Performance Shifts
- DeepSeek-V4.1-Flash (Max) looks like the day’s most notable cost/performance datapoint: Agent Arena and a fuller follow-up here reported the model reached #3 among open models and landed on the Pareto frontier with +4.87% net improvement at roughly $0.06–$0.07 median cost per task. Arena compares that to Hy4 preview at +4.96% / $0.22 and Kimi K3 (Max) at +6.39% / $0.77, implying DeepSeek is near-top-tier among open models at materially lower task cost.
- Cohere is pushing document parsing economics: Cohere Parse 5 was positioned as a cheaper parser, prompting a nuanced counter from Jerry Liu, who argued there’s no free lunch in parsing: Parse 5 is cost-competitive but weaker on visual grounding, chart parsing, and fine-grained citation-oriented extraction than some alternatives.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力