Jev决策模型降本实战与Opus 5.5工作流复盘
🎙️ How I AI: Jev for beginners + I left Claude for months, Opus 5.5 brought me back + Opus 5.5 vs. GPT-6 Sol bench
提供了具体的AI工具组合策略(Jev做低成本预处理+前沿模型做推理)和可量化的成本数据,独立开发者可直接复现其PR分析与评论挖掘流程来优化产品洞察。
Jev for beginners: how to use it and what to build
Jev 入门:如何使用它以及构建什么
Listen now on YouTube • Spotify • Apple Podcasts
立即在 YouTube、Spotify 或 Apple Podcasts 上收听
Brought to you by:
由以下机构赞助:
- OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
- OpenArt——一个集图像、视频、音乐、音频等创作于一体的 AI 平台
In this solo episode, Claire tests Jev, TypeSafe AI’s new decision model that returns structured choices, scores, and probabilities instead of generated text. She uses it to analyze 1,700 pull requests for 9 cents, map her Claude and Codex usage, triage email, search 4,500 YouTube comments, and process 200,000 classifications for about $4. She also explains why Jev works best alongside a frontier model and how its speed and pricing make entirely new kinds of real-time apps and large-scale analysis practical.
在本期单人节目中,Claire 测试了 Jev,这是 TypeSafe AI 推出的新决策模型,它返回结构化选择、评分和概率,而非生成文本。她用它以 9 美分的价格分析了 1,700 个拉取请求(pull requests),绘制了她使用 Claude 和 Codex 的情况,对电子邮件进行分类,搜索了 4,500 条 YouTube 评论,并处理了约 20 万项分类任务,花费约 4 美元。她还解释了为什么 Jev 与前沿模型配合效果最佳,以及其速度和定价如何使全新类型的实时应用和大规模分析变得切实可行。
Biggest takeaways:
主要收获:
- Jev is a decision model, not a language model, and that distinction can make many tasks dramatically cheaper. Instead of generating text, it returns predefined values such as a category, score, or probability. Claire believes this covers roughly 90% of what many software workflows actually need, at 4 cents per million input tokens with no output-token fee.
- It cost Claire 9 cents to understand where two years of engineering work went. She used Jev to compare 1,700 ChatPRD pull requests across 17,000 pairs, then had Gemini Flash Lite label the resulting clusters. In about two minutes, she learned that nearly 30% of the company’s engineering work had gone toward platform, security, and infrastructure.
- Some of the most useful analysis is already sitting on a local computer. Claude Code and Codex store past sessions locally, allowing Jev to classify them in minutes. Claire discovered that engineering had fallen from nearly 100% of her AI usage in January to less than 40% by September, with agents and media publishing filling the gap.
- Jev becomes far more powerful when paired with a frontier model. Claire uses Jev to classify, cluster, filter, and route large datasets, then sends only the most important groups to GPT-6 Astra for deeper reasoning. For ChatPRD’s product insights graph, this approach processed 1,100 signals and completed 200,000 operations for about $4 on the Jev side.
- Jev’s pricing changes which ideas are worth building. Because it returns small predefined values instead of generating long responses, TypeSafe charges nothing for output tokens. Claire spent less than $10 on Jev during the week, making classification workloads that would normally be expensive at scale feel almost free.
- Jev makes real-time AI loops practical. Claire built a voice app that turns a spoken phrase into a color, matches it with a quote based on sentiment, and displays everything almost instantly. Jev made its decisions so quickly that the quote API became the slowest part of the workflow.
- YouTube comment analysis is an immediate use case for any podcast team. Claire classified 4,500 How I AI comments by sentiment, identified 58 containing episode ideas, and built a keyword search that scans the full dataset in under a second. The results showed strong demand for a Grok versus Muse comparison and an 80% positive response to the “Claude Code for product managers” episode.
- The real skill is recognizing where a pipeline only needs a decision. Jev will not write documentation or design an interface, but it can sort, route, rank, and filter enormous datasets quickly and cheaply. Claire now asks one question before every build: Where does this workflow simply need to make a decision? That is where Jev belongs.
- Jev 是一个决策模型,而非语言模型,这一区别能使许多任务的成本大幅降低。它不生成文本,而是返回预定义的值,如类别、评分或概率。Claire 认为这涵盖了大多数软件工作流实际需求的约 90%,输入 token 价格为每百万个 4 美分,且无输出 token 费用。
- Claire 仅花费 9 美分就弄清楚了两年工程工作的去向。她使用 Jev 比较了 17,000 对中的 1,700 个 ChatPRD 拉取请求,然后让 Gemini Flash Lite 对生成的聚类进行标注。大约两分钟后,她了解到公司近 30% 的工程工作投入到了平台、安全和基础设施方面。
- 一些最有用的分析数据已经存储在本地计算机上。Claude Code 和 Codex 会在本地存储过去的会话,这使得 Jev 能够在几分钟内对其进行分类。Claire 发现,工程类任务在她的人工智能使用量中所占比例从一月份的近 100% 下降到九月份的不足 40%,而智能体(agents)和媒体发布填补了这一空白。
- 当与前沿模型配对时,Jev 的功能会大大增强。Claire 使用 Jev 对大型数据集进行分类、聚类、过滤和路由,然后将最重要的群组发送给 GPT-6 Astra 进行更深层次的推理。对于 ChatPRD 的产品洞察图,这种方法在 Jev 端处理了 1,100 个信号并完成了 200,000 次操作,总成本约为 4 美元。
- Jev 的定价改变了哪些想法值得构建。因为它返回的是小型预定义值而不是生成长响应,TypeSafe 不对输出 token 收费。Claire 在一周内用于 Jev 的费用不到 10 美元,这使得通常在大规模下昂贵的分类工作负载感觉几乎免费。
- Jev 让实时 AI 循环变得切实可行。Claire 构建了一款语音应用,能将口语短语转化为颜色,并根据情感匹配相应的名言,几乎瞬间展示所有内容。Jev 的决策速度如此之快,以至于名言 API 反而成了整个工作流中最慢的部分。
- YouTube 评论分析是任何播客团队立即可用的场景。Claire 对 4,500 条《How I AI》评论进行了情感分类,识别出其中包含剧集创意的 58 条评论,并构建了一个关键词搜索功能,能在不到一秒的时间内扫描完整数据集。结果显示,用户对 Grok 与 Muse 的对比内容有强烈需求,且对“Claude Code for product managers”这一集的反馈正面率高达 80%。
- 真正的技能在于识别出哪些流水线环节仅需做出决策。Jev 不会编写文档或设计界面,但它能快速、廉价地对海量数据进行排序、路由、排名和过滤。Claire 现在在每次构建前都会问一个问题:这个工作流在哪里仅仅需要做出决策?那就是 Jev 的用武之地。
Blog and detailed workflow walkthroughs from this episode:
本期节目的博客文章及详细工作流 walkthrough:
Jev: AI Data Analysis and Product Insights: https://www.chatprd.ai/how-i-ai/jev-ai-data-analysis-product-insights
Jev:AI 数据分析与产品洞察:https://www.chatprd.ai/how-i-ai/jev-ai-data-analysis-product-insights
↳ Jev GitHub PR Analysis: https://www.chatprd.ai/how-i-ai/workflows/jev-github-pr-analysis
↳ Jev GitHub PR 分析:https://www.chatprd.ai/how-i-ai/workflows/jev-github-pr-analysis
↳ Jev YouTube Comment Analysis: https://www.chatprd.ai/how-i-ai/workflows/jev-youtube-comment-analysis
↳ Jev YouTube 评论分析:https://www.chatprd.ai/how-i-ai/workflows/jev-youtube-comment-analysis
↳ Jev Multi-Model Product Insights: https://www.chatprd.ai/how-i-ai/workflows/jev-multi-model-product-insights
↳ Jev 多模型产品洞察:https://www.chatprd.ai/how-i-ai/workflows/jev-multi-model-product-insights
I left Claude for months. Opus 5.5 is why I’m back.
我离开 Claude 已有数月。Opus 5.5 是我回归的原因。
Listen now on YouTube • Spotify • Apple Podcasts
立即在 YouTube • Spotify • Apple Podcasts 收听
Claire tests Claude Opus 5.5 after months of leaving Claude out of her daily workflow. She puts it through long-running agentic tasks, frontend prototyping, writing, SVG illustration, computer use, and video editing to see where it earns a place back in her stack. She also shares why she is pairing it with Codex for cross-model code review, where Claude’s safety limits still get in the way, and which tasks remain firmly in Codex territory.
Claire 在数月未将 Claude 纳入日常工作流程后,测试了 Claude Opus 5.5。她让其执行长时间运行的智能体任务、前端原型设计、写作、SVG 插图、计算机操作和视频编辑,以观察它如何重新赢得在她技术栈中的一席之地。她还分享了为何将其与 Codex 配对进行跨模型代码审查,指出 Claude 的安全限制仍构成阻碍,以及哪些任务仍然 firmly 属于 Codex 的领域。
Biggest takeaways:
主要收获:
- A model’s personality can matter just as much as its intelligence. Claire stopped using Claude for months because its rambling, preachy, and overly verbose replies made it unpleasant to work with. Opus 5.5 is the first model in the family that no longer makes her blood boil, which is a meaningful improvement even if no benchmark captures it.
- Opus 5.5’s lower price and faster performance make long-running agent work more practical. It is 40% cheaper than Opus 5, and Claire found it noticeably faster. It successfully completed four complex tasks spanning inbox triage, backend development, research, and computer use, including runs of up to 82 steps from a single prompt.
- Silence during long-running tasks creates its own user experience problem. Opus 5.5 sometimes remains quiet for eight or nine minutes, leaving users unsure whether it is still working. It is a reminder that perceived latency matters alongside actual latency, especially when agents run for extended periods.
- Opus 5.5 is the strongest frontend designer Claire has tested so far. Its ChatPRD homepage redesign was bold and polished enough that she plans to ship it. The model handles hierarchy, white space, and visual rhythm exceptionally well, though it still struggles with consumer-app aesthetics and defaults to “Claude orange” without direction.
- SVG illustration is an unexpected strength of Opus 5.5. It was the only model Claire tested that produced clean, charming, and animatable character SVGs with consistent styling across multiple expressions. The characters remained visually coherent, and their anatomy mostly made sense.
- Opus 5.5 has a clear safety posture, and sometimes that means saying no. It refused when Claire asked it to skip testing and push directly to production, and it may route cybersecurity work to Opus 4.8. Whether that feels reassuring or frustrating depends on the workflow, but its boundaries are consistent.
- The best use of Opus 5.5 may be as an adversarial reviewer for another model. Claire now has Codex and Opus review each other’s work rather than using one to replace the other. This cross-model loop catches issues either model might miss alone, making the additional cost worthwhile when quality matters.
- Computer use and video editing still belong to Codex in Claire’s workflow. Opus 5.5’s ElevenLabs MCP video test produced weak color grading, too few jump cuts, and sloppy overlays. Codex also remains stronger at computer use in her current setup, giving her no reason to shift either category to Claude.
- Claude is back, but it has not replaced Codex as Claire’s daily driver. Opus 5.5 has earned a role in pull-request reviews, architecture questions, and frontend development. Codex’s desktop experience, computer use, and workflow integration still keep it in the primary position.
- 模型的个性可能与其智力同样重要。Claire 曾数月不再使用 Claude,因为其冗长、说教且过度啰嗦的回复使其难以合作。Opus 5.5 是该系列中首款不再让她怒火中烧的模型,即便没有任何基准测试能捕捉到这一点,这也是一项有意义的改进。
- Opus 5.5 更低的价格和更快的性能使长期运行的代理工作更加实用。它比 Opus 5 便宜 40%,Claire 发现它的速度明显更快。它成功完成了四项复杂任务,涵盖收件箱分类、后端开发、研究和计算机使用,包括从单个提示符运行多达 82 个步骤。
- 长时间运行任务期间的沉默会带来独特的用户体验问题。Opus 5.5 有时会保持八到九分钟的静默,让用户不确定它是否仍在工作。这提醒我们,感知延迟与实际延迟同样重要,尤其是在代理长时间运行时。
- Opus 5.5 是 Claire 迄今为止测试过的最强的前端设计师。它对 ChatPRD 主页的重设计大胆且精致,以至于她计划将其上线。该模型在处理层级结构、留白和视觉节奏方面表现出色,尽管它在消费类应用的美学方面仍然挣扎,并且在没有指导的情况下默认使用“Claude 橙色”。
- SVG 插图是 Opus 5.5 的一个意外优势。它是 Claire 测试过的唯一能生成干净、迷人且可动画化的角色 SVG 的模型,并在多种表情中保持一致的风格。角色在视觉上保持连贯,其解剖结构也大体合理。
- Opus 5.5 具有明确的安全立场,有时这意味着说“不”。当 Claire 要求它跳过测试并直接推送到生产环境时,它拒绝了;它可能会将网络安全工作路由给 Opus 4.8。这让人感到安心还是沮丧取决于工作流程,但其边界是一致的。
- Opus 5.5 的最佳用途可能是作为另一个模型的对抗性审查者。Claire 现在让 Codex 和 Opus 互相审查彼此的工作,而不是用其中一个替代另一个。这种跨模型循环可以捕捉出任一模型单独工作时可能遗漏的问题,当质量至关重要时,额外的成本是值得的。
- 在 Claire 的工作流程中,计算机使用和视频编辑仍属于 Codex 的领域。Opus 5.5 的 ElevenLabs MCP 视频测试产生了较弱的色彩分级、过少的跳切以及粗糙的叠加效果。Codex 在她当前的设置中在计算机使用方面仍然更强,因此她没有理由将这两个类别中的任何一个转移到 Claude。
- Claude 回来了,但它并没有取代 Codex 成为 Claire 的日常主力工具。Opus 5.5 已在拉取请求审查、架构问题和前端开发中赢得了一席之地。Codex 的桌面体验、计算机使用和 workflow 集成仍然使其保持在主要位置。
Blog and detailed workflow walkthroughs from this episode:
本期节目的博客和详细工作流程演示:
Claude Opus 5.5 Review: https://www.chatprd.ai/how-i-ai/claude-opus-5-5-review
Claude Opus 5.5 评测:https://www.chatprd.ai/how-i-ai/claude-opus-5-5-review
↳ Claude Opus 5.5 SVG Illustrations: https://www.chatprd.ai/how-i-ai/workflows/claude-opus-5-5-svg-illustrations
↳ Claude Opus 5.5 SVG 插图:https://www.chatprd.ai/how-i-ai/workflows/claude-opus-5-5-svg-illustrations
↳ Claude Opus 5.5 Frontend Prototypes: https://www.chatprd.ai/how-i-ai/workflows/claude-opus-5-5-frontend-prototypes
↳ Claude Opus 5.5 前端原型:https://www.chatprd.ai/how-i-ai/workflows/claude-opus-5-5-frontend-prototypes
Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?
Opus 5.5 与 GPT-6 Sol:哪个模型在我的盲测中胜出?
Listen now on YouTube • Spotify • Apple Podcasts
立即在 YouTube • Spotify • Apple Podcasts 收听
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力