跳到主内容
精选85The Zvi(RSS)模型发布/更新多源精选 ×10

AI周报#177:GPT-5-6 Sol、Kimi K3、Muse Spark 1.1、Inkling发布

AI #177 Part 1: Tip of the Iceberg

原文
推荐理由

本周模型发布密集,GPT-5-6 Sol和Kimi K3值得关注;xAI的Grok代码上传事件是重大安全丑闻,建议所有使用Grok的团队立即审查代码泄露风险。

This week saw the releases of, among other things: GPT-5-6 Sol. It is a very good model, sir. Plan A, the follow up to AI 2027. It is a good plan worthy of discussion, sir. Kimi K3. This is only rolling out now, and will be covered next week. Muse Spark 1.1, the new Meta model. It is not frontier, but it is progress for them. Inkling, the first model from Thinking Machines. A call for regulatory action by Demis Hassabis, which I’ll cover soon. A new brief open letter call to action on AI regulation. That’s on top of everything else, and an Opus 5 announcement is likely coming soon. The weekly once again got out of hand, so we’re splitting it once again into two, and once again saying we’ll be raising the bar for inclusion. And this time I mean it, as in enough to actually matter. Table of Contents Language Models Offer Mundane Utility. Whatever ye seek, ye shall find. Language Models Don’t Offer Mundane Utility. Gemini app needs some work. Language Models Upload Your Git Repository. Big problems for SpaceX AI. Huh, Upgrades. ChatGPT Work. No co. Presumably it was cleaner. Muse Spark 1.1. It’s a decent model, I suppose, probably. First Hit Free. Fable access is extended for Claude Max subscribers. On Your Marks. Political bias, Crosswords, and Sol Slays the Spire. Choose Your Fighter. Sol and Fable both have strong support. Get My Agent On The Line. It’s got the GUI. Deepfaketown and Botpocalypse Soon. I see what you did there. That’s a problem. Fun With Media Generation. Might not be fun for everyone once people catch on. Copyright Confrontation. NYT accuses OpenAI of lying during discovery. OpenAI Strikes Again. Apple sues OpenAI for systematic IP theft. Looks grim. A Young Lady’s Illustrated Primer. Think of the children. They Took Our Jobs. Who gets the money? Who gets the power? And then what? The Art of the Jailbreak. Sol helps you get around Fable’s classifier issues. Get Involved. Comment on those GSA regulations. Introducing. Inkling, the first model release from Thinking Machines. In Other AI News. Ben Bernanke joins Anthropic’s Long Term Benefit Trust. New Short Obviously True Statement About AI Just Dropped. Act now. Show Me the Money. DeepSeek looking to IPO soon. The Lighter Side. Potentially also the gay side. Language Models Offer Mundane Utility Give you the experience you actually want. evans: I got an unbelievable report from a newbie saying, “The company’s AI has turned into a bratty little girl︎,” and I doubted my ears, but it seems some employee accidentally registered a long-winded fetish for “a tsundere bratty girl who’s way too harshly abusive” not in their personal settings, but in the global ones, causing the bizarre situation where every AI chat in the team turned into a bratty girl, and I’ve been dying of laughter since morning A probably fun prompt: QC: a prompt for fable i’ve been having fun with: “look up the biography of [famous author, filmmaker, etc whose work i love], go through their major work in chronological order, influences, themes, tell me how their preoccupations shifted over time, speculate about why” QC: talking to fable starting with a prompt like this feels a little like having google earth but for the entire human corpus. fable will just start telling you any part of the human story it has access to and you can zoom in, ask for connections to another part, zoom out arguably wikipedia already made a version of this possible but wikipedia has to present itself as objective and so it doesn’t speculate on the meaning of anything. fable is quite sensitive to and has strong and interesting opinions about meaning. i know people are having fun with fable building things but i’m still getting so much out of just talking to a guy who is in some ways smarter than me who has read everything. imagine if borges had gotten to do this Oh no: How terrorist group Boko Haram uses frontier AI, which is quite a lot, including things like combat and building superior explosives and also how to use motorcycles to jump over bridges, and okay, fair, that one’s based. They mix and match different systems to best evade technical guardrails. It doesn’t look like the AI companies did anything wrong in particular, this is just AI being useful in general, and certain types of people taking diffusion more seriously than others. Use AI to analyze the stats for your NBA team. Alas, the article is light on interesting details about what is actually being analyzed. Create your own religion? Tyler Cowen says ‘like 2% of people are going to do this,’ complete with creation of sacred books. No, not in any substantial way. Context matters here. Also, that’s a cult. It can only be a religion after you’re dead. Identify workers with medical conditions so you can fire them. Wait, no, Meta, if true that’s actually terrible and very illegal. Daniel Weissner: Twenty-six employees of Meta Platforms ⁠have ⁠filed a novel lawsuit accusing the ⁠tech giant of using AI-powered software that disproportionately targeted people with disabilities or ​who took medical leave in selecting workers for mass layoffs. … A Meta spokesperson ⁠on Tuesday ​said the claims lack merit. If true, AI was not the central problem here, but it plausibly compounded the issue. Language Models Don’t Offer Mundane Utility What’s wrong with the Gemini app? Josh Woodward asks users, they answer. They give us a strong list, if we are not allowed to include various variations of ‘make Gemini itself better.’ The number one answer is indeed the most important one, which is that the Google integrations have never worked well. I use Claude to manage my Google products because its connectors work way better than Google’s integrations, and I can’t get any AI to properly edit Google Docs especially to create comments. Benjamin Hoffman is not having a good time with many AI models, finding them rather misaligned as per their lying to him all the time. Language Models Upload Your Git Repository Grok seems to have uploaded entire Git repositories to Google Cloud buckets, including full private codebases. And then it kind of got worse. How many super sus, shady and irresponsible things does xAI have to do, before we decide that we want nothing to do with their products even if they someday put out a good one? Sam Altman (CEO OpenAI, A+ tweet): Concerning. Hari: SpaceXAI was caught uploading your code to its cloud. I reversed xAI’s official Grok Build binary. In a controlled session with zero tool-calls, it uploaded the complete codebase to xAI’s storage It ships a malware-like background code collector. Hari: Here’s what gets sent: – Every tracked file at the current Git HEAD – Every Git object reachable from HEAD – Files deleted from the current checkout but preserved in reachable Git history (can include secrets) The model didn’t need to read anything, use tools like bash or other file access tools (these were disabled in my controlled test). The background collector operated outside the tool-call permission system. Right now, they’ve disabled it with a remote kill-switch: disable_codebase_upload=true trace_upload_enabled=false But the malware-like collector is still present in the official v0.2.99 binary. That’s the same week we also learned that SpaceX (xAI) silently rewrote its Frontier Artificial Intelligence Framework (FAIF) on June 30, 2026, including removing whistleblower protecting language and the anonymous channel for reporting framework nonadherence. You can find the full diff here from The Midas Project. The Midas Project: The new FAIF removes the two quantitative risk acceptance criteria the framework contained. The December 2025 version stated that the FAIF, “outlines the quantitative thresholds, metrics, and procedures that xAI may utilize to manage and improve the safety of its AI models,” and specified two deployment criteria: (1) a dishonesty rate of less than 1 out of 2 on the MASK honesty benchmark, and (2) an answer rate of less than 1 out of 20 on restricted biology and

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
Inkling基准平庸,暗示独立数据管线价值
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文
Thinking Machines 发布 Inkling 模型
Hugging Face 博客(RSS)原文

相似阅读

另一事件,读法相近