Grok Bot、Cursor Origin、Grok 4.6 实测
🎙️ How I AI: Grok Bot + Grok 4.6—what’s great (and what’s still hype) & Lessons from spending $20,000 on Devin in one month
Grok Bot + Grok 4.6—what’s great (and what’s still hype)
Grok Bot + Grok 4.6——哪些令人惊艳(哪些仍是炒作)
Listen now on YouTube • Spotify • Apple Podcasts
现在在 YouTube • Spotify • Apple Podcasts 上收听
Brought to you by:
由以下赞助:
Bolt.new—Turn your idea into a real product
Bolt.new——将你的想法变成真正的产品
Jira AI SDLC—Get your tokens’ worth with Jira
Jira AI SDLC——用 Jira 让你的代币物有所值
In this solo episode, Claire tests Grok Bot, Cursor Origin, and Grok 4.6 to figure out what’s genuinely useful and what’s still mostly hype. She shares the Grok Bot feature that immediately won her over, why she isn’t ready to replace GitHub with Origin, and how Grok 4.6 performed against GPT-5.6, Claude Sonnet 5, and Opus 5 in her own blind evaluations.
在这期单人节目中,Claire 测试了 Grok Bot、Cursor Origin 和 Grok 4.6,以分辨哪些真正实用,哪些仍主要是炒作。她分享了立即赢得她青睐的 Grok Bot 功能,为什么她还没准备好用 Origin 取代 GitHub,以及 Grok 4.6 在她自己的盲评中与 GPT-5.6、Claude Sonnet 5 和 Opus 5 相比表现如何。
Biggest takeaways:
最重要的收获:
- Grok Bot’s multi-account connectors solve a problem every other agent platform seems to ignore. Most platforms assume each person has one Gmail account and one Slack workspace. Claire has four email addresses and seven Slack workspaces. Grok Bot lets her connect all of them to a single bot, which made it genuinely useful from day one in a way that Codex and Claude still have not matched.
- Grok Bot’s simplicity is both its greatest strength and its biggest limitation. Setup is fast, the iMessage-style interface is clean, and the built-in plugins actually work. But people who enjoy customizing their agents, choosing models, shaping personalities, and tinkering with every detail may find it almost too polished. Claire loves her OpenClaw agents partly because they are chaotic and high-maintenance. So far, Grok Bot has not given her much to wrestle with.
- Cursor Origin is a compelling vision that is not quite ready for prime time. An agent-native alternative to GitHub makes a lot of sense, especially one where Bugbot, Cursor, and the entire pull request workflow are designed around how coding agents actually work. Today, though, Origin still feels like a more attractive version of GitHub with fewer features. Teams that rely heavily on GitHub Actions, code owners, and existing automations will need a much stronger reason to migrate.
- Grok 4.6 is a genuine frontier-model competitor. That conclusion did not come from someone else’s leaderboard. Claire runs her evaluations blind, grades the outputs herself, and gives her own judgment 70% of the final weight. Grok 4.6 finished alongside GPT-5.6 Sol at the top of the Claire Index, ahead of both Sonnet 5 and Opus 5.
- For sharp, enjoyable agent conversations, Sonnet 5 is still the model to beat. When Claire wants an OpenClaw agent that is concise, responsive, and fun to talk with, Sonnet 5 continues to win. It has a conversational rhythm that feels more like working with a strong collaborator than issuing commands to a tool. Grok 4.6 does not yet compete in that category.
- Cursor and xAI are assembling a surprisingly coherent enterprise stack. Grok Bot serves knowledge workers, Origin handles code hosting, Grok 4.6 provides a capable default model, and the Cursor IDE already sits at the center of many developers’ workflows. None of the pieces is perfect on its own, but together they are beginning to look like a credible enterprise platform. Large companies often prefer one vendor that can own the entire experience, and Cursor is increasingly positioned to become that vendor.
- Grok Bot 的多账户连接器解决了其他所有代理平台似乎都忽略的问题。大多数平台假设每个人只有一个 Gmail 账户和一个 Slack 工作区。Claire 有四个电子邮件地址和七个 Slack 工作区。Grok Bot 让她能将所有这些连接到一个机器人上,这使得它从第一天起就真正实用,而 Codex 和 Claude 至今仍未能做到这一点。
- Grok Bot 的简洁性既是其最大优势,也是其最大局限。设置快速,iMessage 风格的界面干净整洁,内置插件也确实有效。但喜欢自定义代理、选择模型、塑造个性并调整每个细节的人可能会觉得它过于精致。Claire 喜欢她的 OpenClaw 代理,部分原因是它们混乱且需要高维护。到目前为止,Grok Bot 并没有给她太多需要费力应对的地方。
- Cursor Origin 是一个引人注目的愿景,但尚未完全准备好迎接黄金时段。一个以代理为中心的 GitHub 替代品很有意义,尤其是像 Bugbot、Cursor 和整个拉取请求工作流都围绕编码代理的实际工作方式设计的那种。然而,目前 Origin 感觉更像是 GitHub 的一个更具吸引力的版本,但功能更少。严重依赖 GitHub Actions、代码所有者和现有自动化的团队将需要更充分的理由来迁移。
- Grok 4.6 是一个真正的领先模型竞争者。这个结论并非来自别人的排行榜。Claire 进行盲评,自己给输出打分,并给予自己的判断 70% 的最终权重。Grok 4.6 与 GPT-5.6 Sol 一起位居 Claire 指数榜首,领先于 Sonnet 5 和 Opus 5。
- 对于敏锐、愉快的代理对话,Sonnet 5 仍然是难以超越的模型。当 Claire 想要一个简洁、响应迅速且对话有趣的 OpenClaw 代理时,Sonnet 5 依然胜出。它的对话节奏更像是在与一位强大的协作者合作,而非向工具下达命令。Grok 4.6 在这一类别中尚无法匹敌。
- Cursor 和 xAI 正在构建一个出奇一致的企业级技术栈。Grok Bot 服务知识工作者,Origin 处理代码托管,Grok 4.6 提供了一款能力出众的默认模型,而 Cursor IDE 已经处于许多开发者工作流程的核心位置。这些组件单独来看并非完美,但组合在一起,它们开始看起来像一个可信的企业平台。大型公司往往倾向于选择能提供整体体验的单一供应商,而 Cursor 正日益成为这样的供应商。
Blog and detailed workflow walkthroughs from this episode:
本集的博客和详细工作流程指南:
My Hands-On Review of GrokBot, Cursor Origin, and the Grok 4.6 Model: https://www.chatprd.ai/how-i-ai/how-i-ai-my-hands-on-review-of-grokbot-cursor-origin-and-the-grok-46-model
我对 GrokBot、Cursor Origin 和 Grok 4.6 模型的亲身体验评测:https://www.chatprd.ai/how-i-ai/how-i-ai-my-hands-on-review-of-grokbot-cursor-origin-and-the-grok-46-model
↳ How to Automate Knowledge Work Across Multiple Accounts with GrokBot: https://www.chatprd.ai/how-i-ai/workflows/how-to-automate-knowledge-work-across-multiple-accounts-with-grokbot
↳ 如何使用 GrokBot 跨多个账户自动化知识工作:https://www.chatprd.ai/how-i-ai/workflows/how-to-automate-knowledge-work-across-multiple-accounts-with-grokbot
I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)
我一个月在 Devin 上花了 2 万美元。以下是我的心得 | Ryan Carson(独立创始人)
Listen now on YouTube • Spotify • Apple Podcasts
现在在 YouTube • Spotify • Apple Podcasts 上收听
Brought to you by:
由以下赞助:
- WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more
- Jira AI SDLC—Get your tokens’ worth with Jira
- WorkOS——通过 SSO、SCIM、RBAC 等,让您的应用达到企业级标准
- Jira AI SDLC——让您的令牌物有所值,尽在 Jira
Ryan Carson is a five-time founder and the solo founder of Untangle, a B2B SaaS platform for family law firms. In this episode, he breaks down how he manages up to 15 AI agents at once, ships as many as 40 pull requests a day, and uses Devin, Codex, and Claude Code to handle everything from engineering and QA to customer success and investor updates. He also shares why more AI output doesn’t necessarily lead to a better product, how a handwritten priority list keeps his agents focused, and why talking to one real customer changed the direction of his entire company.
Ryan Carson 是一位五次创业的创始人,也是 Untangle 的独立创始人,Untangle 是一个面向家庭律师事务所的 B2B SaaS 平台。在本集中,他详细讲述了如何同时管理多达 15 个 AI 代理,每天提交多达 40 个拉取请求,并利用 Devin、Codex 和 Claude Code 处理从工程、质量保证到客户成功和投资者更新的一切事务。他还分享了为什么更多的 AI 输出并不一定带来更好的产品,手写的优先级列表如何让他的代理保持专注,以及为什么与一位真实客户的交谈改变了他整个公司的方向。
Biggest takeaways:
最重要的收获:
- The most important skill for a solo founder may be managing agents, not writing code. Ryan runs 10 to 15 Devin threads at once, organized into P0, P1, P2, and Bugs folders. He treats each thread the way a good manager treats a direct report: give it a clear goal, set the right priority, and avoid unnecessary hand-holding. That organizational discipline is what separates founders who gain real leverage from AI from those who simply accumulate open tabs.
- A piece of paper can still be the best tool for protecting attention. Despite working across eight screens, Ryan keeps a handwritten list of his weekly priorities beside him. The simplicity is intentional. When agents are constantly generating updates, questions, and decisions, that physical list keeps him anchored to the three things that matter most before he gets pulled into whatever Devin surfaced overnight.
- Ryan built an AI playbook that functions like a customer success team. His Watchdog workflow reviews every law firm account, pulls recent activity and errors from Sentry and internal logs, identifies the three most important problems, and checks whether a recent pull request has already fixed each one. Everything appears in a single Devin thread. Whenever Ryan feels the familiar anxiety of not knowing what is happening across his customer base, he runs Watchdog and quickly gets back up to speed.
- Producing more with AI does not automatically lead to a better product. Both Ryan and Claire are skeptical of letting agents work without constraints overnight. Frontier models can generate an enormous amount of output, but they do not know what customers actually need. Product ideas, priorities, and market judgment still have to come from a human who talks to users. Ryan found product-market fit for Untangle not by shipping more code but by landing one meeting with family law attorney Renee Bauer and listening carefully.
- Cloud agents may handle most engineering work, but Codex still shines when the frontend requires close attention. Ryan uses Devin for cloud-based work across bugs, pull requests, investor updates, and customer triage. He turns to Codex when a feature is visually complex and he needs to stay close to the implementation. The distinction is not about loyalty to a particular tool. It comes down to latency, browser access, and the ability to inspect and refine the interface in real time.
- Coding agents can operate far beyond the codebase. Claire uses Devin for deal desk, custom quotes, and customer triage. Ryan uses it to prepare investor updates through a reusable skill. Both approach these agents with a broader question: What could someone accomplish if they understood the entire codebase and could write software to solve almost any problem? That framing reveals business uses that would never emerge from treating an agent as a simple coding assistant.
- The best engineering interview may be a screen recording of the actual work. Ryan asks candidates to record themselves building a feature inside an existing application, with the entire screen visible and no introductory meeting required. In the next stage, he gives them access to Devin and reviews the replay of how they worked with it. This shows him how candidates think, build, and manage an agent under realistic conditions, which he finds far more revealing than a typical behavioral interview.
- 对于独立创始人来说,最重要的技能可能是管理智能体,而不是编写代码。Ryan 同时运行着 10 到 15 个 Devin 线程,并组织成 P0、P1、P2 和 Bugs 文件夹。他对待每个线程的方式,就像一位优秀经理对待直接下属一样:设定明确目标,确定正确优先级,避免不必要的过度干预。这种组织纪律,正是那些真正从 AI 中获得杠杆效应的创始人与那些只是积累一堆打开标签页的创始人之间的分水岭。
- 一张纸仍然可能是保护注意力的最佳工具。尽管在八块屏幕上工作,Ryan 手边仍放着一份手写的每周优先事项清单。这种简单是有意为之。当智能体不断生成更新、问题和决策时,这份实体清单能让他锚定在最重要的三件事上,而不会被他被 Devin 一夜之间冒出来的内容拉走。
- Ryan 构建了一套 AI 行动手册,其功能相当于一个客户成功团队。他的 Watchdog 工作流会审查每个律师事务所账户,从 Sentry 和内部日志中提取近期活动和错误,识别出三个最重要的问题,并检查最近的拉取请求是否已修复每个问题。所有内容都出现在一个 Devin 线程中。每当 Ryan 感到那种熟悉的焦虑——不知道客户群中正在发生什么时,他就会运行 Watchdog,迅速重新掌握情况。
- 用 AI 产出更多并不自动意味着更好的产品。Ryan 和 Claire 都对让智能体在无约束条件下过夜工作持怀疑态度。前沿模型可以生成大量输出,但它们不知道客户真正需要什么。产品创意、优先级和市场判断仍然必须来自与用户交流的人。Ryan 为 Untangle 找到产品市场契合点,不是通过发布更多代码,而是通过约见家庭法律师 Renee Bauer 并仔细倾听。
- 云端智能体可能处理大部分工程工作,但当前端需要密切注意时,Codex 仍然表现出色。Ryan 使用 Devin 处理基于云的各项工作,包括 bug、拉取请求、投资者更新和客户分类。当某个功能在视觉上复杂且他需要贴近实现时,他会转向 Codex。这种区分并非出于对特定工具的忠诚,而是归结于延迟、浏览器访问以及实时检查和优化界面的能力。
- 编码代理的作用范围远不止代码库。克莱尔利用Devin处理交易台、定制报价和客户分流。瑞安则通过可复用的技能,用它来准备投资者更新。两人都以更广阔的视角看待这些代理:如果某人能理解整个代码库,并能编写软件解决几乎任何问题,那他能达成什么?这种框架揭示了那些仅将代理视为简单编码助手永远不会浮现的业务用途。
- 最好的工程面试可能是一段实际工作的屏幕录制。瑞安要求候选人在现有应用中录制自己构建功能的过程,全程屏幕可见,无需初步会议。在下一阶段,他给予他们Devin的访问权限,并审查他们如何与之协作的回放。这让他看到候选人在现实条件下如何思考、构建和管理代理,他认为这比典型的行为面试更具启发性。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力