跳到主内容
@wquguru
精选75DAIR.AI(RSS)行业动态

AI Agents Weekly:Agent插件标准、Qwen3.8-Max等

🤖 AI Agents Weekly: Agent Plugins Standard, Qwen3.8-Max, Meta Muse Code, Prime Agent, LFM2.5-2.6B, Qwen-CUA, Harness Evolution Papers, and More

原文
发到 X

In today’s issue:

  • Six labs ship Agent Plugins standard
  • Alibaba launches Qwen3.8-Max at 2.4T
  • Meta ships Muse Code terminal agent
  • Prime Intellect open-sources Prime Agent
  • Liquid AI ships on-device agent model
  • OpenAI model solves 10 open problems
  • OpenRouter launches Ori Harness
  • Firecrawl open-sources anydoc parser
  • Cursor open-sources MoE megakernel
  • Mistral drops Shieldstral safety model
  • Cloudflare defines agent dev lifecycle
  • Qwen-CUA hits 86.2 on OSWorld
  • HarnessCompass evolves agent harnesses
  • Harness-R1 learns to patch runtimes
  • AutoCompact learns when to compact
  • Memory consolidation breaks authority

And all the top AI dev news, papers, and tools.

Top Stories

Agent Plugins Standard Launches

AWS, Cursor, GitHub, Microsoft, OpenAI, and Vercel jointly announced Agent Plugins, an open packaging standard for extending agents across clients.

  • One format, many clients: A plugin is a folder with a plugin.json manifest that bundles Agent Skills and MCP server configs, loadable by any compatible client.
  • Launch support: ChatGPT, Codex, GitHub Copilot, VS Code, Cursor, and Kiro all read the same package on day one.
  • Why it matters: Until now every product used its own folder layout and install flow, forcing builders to repackage the same extension per platform.
  • Governance: The spec is public with a technical steering committee spanning Amazon, Cursor, Microsoft, OpenAI, and Vercel, with proposals handled in the open.

Spec | Blog

Alibaba Launches Qwen3.8-Max

Alibaba released Qwen3.8-Max, its largest flagship model, aimed squarely at long-horizon coding and professional work.

  • Scale: 2.4T total parameters with 95B active and a 1-million-token context window, built on the Qwen3.5 architecture.
  • Agentic benchmarks: 86.6 on TerminalBench 2.1, 86.1 on OSWorld-Verified, 93.0 on PaperBench, and 74.8 on CoWorkBench.
  • Long-horizon evidence: A roughly 16-day autonomous run on the oh-my-cli project produced 265 commits, 127 pull requests, and 151 issues, and the model beat 458 of 526 human teams in a multimodal dialogue challenge.
  • Availability: Live on QwenCloud with OpenAI and Anthropic protocol support plus Claude Code, Codex, Qoder CLI, and OpenClaw integrations, with open weights promised alongside Qwen3.8-27B.

Blog

Read more

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件