精选75DAIR.AI(RSS)行业动态多源精选 ×8
DeepSeek 开源 Agent Harness,发布 V4-Pro
🤖 AI Agents Weekly: DeepSeek Harness, DeepSeek-V4-Pro, Grok Bot, GLM-5.3, Gemini 3.7 Flash, Muse Glimmer, Harness Evolution Papers, and More
In today’s issue:
今日要闻:
- DeepSeek open-sources its agent harness
- DeepSeek-V4-Pro ships agent upgrades
- xAI launches Grok Bot teammates
- Z.ai drops GLM-5.3 for coding
- Gemini 3.7 Flash halves coding cost
- Meta open-sources Muse Glimmer
- Grok 4.6 hits frontier at half price
- Zed launches Delta for agent teams
- Evo-Bench measures harness evolution
- Study finds 91.8% of skills defective
- DeepSeek 开源其智能体框架
- DeepSeek-V4-Pro 推出智能体升级
- xAI 推出 Grok Bot 队友
- Z.ai 发布 GLM-5.3 用于编程
- Gemini 3.7 Flash 将编程成本减半
- Meta 开源 Muse Glimmer
- Grok 4.6 以半价达到前沿水平
- Zed 推出面向智能体团队的 Delta
- Evo-Bench 衡量框架演进
- 研究发现 91.8% 的技能存在缺陷
And all the top AI dev news, papers, and tools.
以及所有顶级 AI 开发新闻、论文和工具。
Top Stories
头条新闻
DeepSeek Open-Sources Its Agent Harness
DeepSeek 开源其智能体框架
DeepSeek released DeepSeek Harness v0.1 as a developer preview, open-sourcing the codebase under MIT and opening it to anyone building agent harnesses.
DeepSeek 发布了 DeepSeek Harness v0.1 作为开发者预览版,以 MIT 许可证开源代码库,并向所有构建智能体框架的人开放。
- Everything is a plugin: The harness is built on the Cordis meta-framework, a kernel that mounts, unmounts, and resolves dependencies for models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and UI as independent plugins.
- Append-only session log: Everything the model sees is recorded, so sessions can be resumed, forked, searched, and replayed rather than reconstructed from chat history.
- Four runtime modes: Standard ships the full toolset, Code orchestrates operations through TypeScript, Minimal strips down for benchmark runs, and Creator is for building custom presets.
- Install path: Runs via npx @deepseek-ai/dsh web or from source, and the repo has already cleared 93,000 stars.
- 一切皆插件:该框架基于 Cordis 元框架构建,这是一个内核,可挂载、卸载并解析模型、工具、技能、会话、沙箱、存储、循环、调度和 UI 作为独立插件的依赖关系。
- 仅追加会话日志:模型看到的所有内容都会被记录,因此会话可以恢复、分叉、搜索和重放,而不是从聊天历史中重建。
- 四种运行时模式:Standard 提供完整工具集,Code 通过 TypeScript 编排操作,Minimal 为基准测试精简,Creator 用于构建自定义预设。
- 安装路径:通过 npx @deepseek-ai/dsh web 或从源码运行,该仓库已获得超过 93,000 颗星。
Blog | GitHub
博客 | GitHub
DeepSeek Launches V4-Pro
DeepSeek 推出 V4-Pro
DeepSeek shipped V4-Pro-0813, a general availability release centered almost entirely on agent workloads.
DeepSeek 发布了 V4-Pro-0813,这是一个几乎完全围绕智能体工作负载的通用可用性版本。
- Agentic benchmarks: 87.9 on Terminal Bench 2.1, 62.7 on DeepSWE, 74.1 on Toolathlon-Verified, 83.3 on CyberGym, and 31.8 on public AutomationBench, tested through DeepSeek Harness in minimal mode.
- Flexible reasoning effort: Low, high, and max tiers across V4-Pro and V4-Flash let you dial spend per task instead of paying reasoning cost on trivial calls.
- Native Responses API: Ships OpenAI Responses API support with one-click Codex setup, and model names stay unchanged so existing integrations keep working.
- Peak and off-peak pricing: New API rates take effect August 16, with off-peak rates 50% below peak for schedulable batch and agent workloads.
- 智能体基准测试:在 Terminal Bench 2.1 上得分 87.9,DeepSWE 上 62.7,Toolathlon-Verified 上 74.1,CyberGym 上 83.3,公共 AutomationBench 上 31.8,通过 DeepSeek Harness 以最小模式测试。
- 灵活的推理努力:V4-Pro 和 V4-Flash 上的低、高、最大级别让您可以根据任务调整支出,而不是为琐碎调用支付推理成本。
- 原生响应API:支持OpenAI响应API,一键设置Codex,模型名称保持不变,现有集成继续工作。
- 高峰和低谷定价:新的API费率于8月16日生效,低谷费率比高峰低50%,适用于可调度的批处理和代理工作负载。
Blog
博客
Read more
阅读更多
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力