跳到主内容
精选75DAIR.AI(RSS)行业动态多源精选 ×8

DeepSeek 开源 Agent Harness,发布 V4-Pro

🤖 AI Agents Weekly: DeepSeek Harness, DeepSeek-V4-Pro, Grok Bot, GLM-5.3, Gemini 3.7 Flash, Muse Glimmer, Harness Evolution Papers, and More

原文

In today’s issue:

今日要闻:

  • DeepSeek open-sources its agent harness
  • DeepSeek-V4-Pro ships agent upgrades
  • xAI launches Grok Bot teammates
  • Z.ai drops GLM-5.3 for coding
  • Gemini 3.7 Flash halves coding cost
  • Meta open-sources Muse Glimmer
  • Grok 4.6 hits frontier at half price
  • Zed launches Delta for agent teams
  • Evo-Bench measures harness evolution
  • Study finds 91.8% of skills defective
  • DeepSeek 开源其智能体框架
  • DeepSeek-V4-Pro 推出智能体升级
  • xAI 推出 Grok Bot 队友
  • Z.ai 发布 GLM-5.3 用于编程
  • Gemini 3.7 Flash 将编程成本减半
  • Meta 开源 Muse Glimmer
  • Grok 4.6 以半价达到前沿水平
  • Zed 推出面向智能体团队的 Delta
  • Evo-Bench 衡量框架演进
  • 研究发现 91.8% 的技能存在缺陷

And all the top AI dev news, papers, and tools.

以及所有顶级 AI 开发新闻、论文和工具。

Top Stories

头条新闻

DeepSeek Open-Sources Its Agent Harness

DeepSeek 开源其智能体框架

DeepSeek released DeepSeek Harness v0.1 as a developer preview, open-sourcing the codebase under MIT and opening it to anyone building agent harnesses.

DeepSeek 发布了 DeepSeek Harness v0.1 作为开发者预览版,以 MIT 许可证开源代码库,并向所有构建智能体框架的人开放。

  • Everything is a plugin: The harness is built on the Cordis meta-framework, a kernel that mounts, unmounts, and resolves dependencies for models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and UI as independent plugins.
  • Append-only session log: Everything the model sees is recorded, so sessions can be resumed, forked, searched, and replayed rather than reconstructed from chat history.
  • Four runtime modes: Standard ships the full toolset, Code orchestrates operations through TypeScript, Minimal strips down for benchmark runs, and Creator is for building custom presets.
  • Install path: Runs via npx @deepseek-ai/dsh web or from source, and the repo has already cleared 93,000 stars.
  • 一切皆插件:该框架基于 Cordis 元框架构建,这是一个内核,可挂载、卸载并解析模型、工具、技能、会话、沙箱、存储、循环、调度和 UI 作为独立插件的依赖关系。
  • 仅追加会话日志:模型看到的所有内容都会被记录,因此会话可以恢复、分叉、搜索和重放,而不是从聊天历史中重建。
  • 四种运行时模式:Standard 提供完整工具集,Code 通过 TypeScript 编排操作,Minimal 为基准测试精简,Creator 用于构建自定义预设。
  • 安装路径:通过 npx @deepseek-ai/dsh web 或从源码运行,该仓库已获得超过 93,000 颗星。

Blog | GitHub

博客 | GitHub

DeepSeek Launches V4-Pro

DeepSeek 推出 V4-Pro

DeepSeek shipped V4-Pro-0813, a general availability release centered almost entirely on agent workloads.

DeepSeek 发布了 V4-Pro-0813,这是一个几乎完全围绕智能体工作负载的通用可用性版本。

  • Agentic benchmarks: 87.9 on Terminal Bench 2.1, 62.7 on DeepSWE, 74.1 on Toolathlon-Verified, 83.3 on CyberGym, and 31.8 on public AutomationBench, tested through DeepSeek Harness in minimal mode.
  • Flexible reasoning effort: Low, high, and max tiers across V4-Pro and V4-Flash let you dial spend per task instead of paying reasoning cost on trivial calls.
  • Native Responses API: Ships OpenAI Responses API support with one-click Codex setup, and model names stay unchanged so existing integrations keep working.
  • Peak and off-peak pricing: New API rates take effect August 16, with off-peak rates 50% below peak for schedulable batch and agent workloads.
  • 智能体基准测试:在 Terminal Bench 2.1 上得分 87.9,DeepSWE 上 62.7,Toolathlon-Verified 上 74.1,CyberGym 上 83.3,公共 AutomationBench 上 31.8,通过 DeepSeek Harness 以最小模式测试。
  • 灵活的推理努力:V4-Pro 和 V4-Flash 上的低、高、最大级别让您可以根据任务调整支出,而不是为琐碎调用支付推理成本。
  • 原生响应API:支持OpenAI响应API,一键设置Codex,模型名称保持不变,现有集成继续工作。
  • 高峰和低谷定价:新的API费率于8月16日生效,低谷费率比高峰低50%,适用于可调度的批处理和代理工作负载。

Blog

博客

Read more

阅读更多

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近