AI #174:GLM-5.2 成最强开源模型,Claude Tag 接入 Slack
AI #174: You’re It
Fable remains in limbo, with renewed hope that we will get it back soon (45% by tomorrow, 69% by July 1, nice.) The full capabilities post is now available. Alex Bores unfortunately lost narrowly in NY-12, and will not be heading to Congress. There are also plenty of other stories to cover. Some highlights: GLM-5.2 is the new best open model, although it is expensive for its class. It will have its uses, potentially for agents you need to run fully locally or privately, but often it won’t be the right fit. Claude Tag is a new system for having Claude join your Slack, and if you @ him then he will spin up an instance to do the coding work. Dean Ball is joining OpenAI to work on policy. We don’t see eye to eye on everything, but this is a huge upgrade over their existing alternatives. The debate over the MidJourney scanner continues. Table of Contents Language Models Offer Mundane Utility. You know what it is for. Language Models Don’t Offer Mundane Utility. Hiring French Qwants. Huh, Upgrades. Claude Code supports artifacts. On Your Marks. Refine reviews economics papers. Deepfaketown and Botpocalypse Soon. We don’t need Pangram to spot it. Fun With Media Generation. This AI production seemed less obvious. Cyber Lack of Security. Five eyes, yet no humans involved. Overcoming Bias. Bias as trust in those with a political affiliation. A Young Lady’s Illustrated Primer. A lot of newfound failure. They Took Our Jobs. Jeff Bezos tells quite the whopper. Get Involved. FLI AI for epistemics prize, Johns Hopkins fellowship. Introducing. Are you in the weights? Claude Tag. He’s in your Slack, ready to spin up an instance. In Other AI News. Various ways to play to win the games. More On GLM-5.2. The claimed niche is ‘open top level agent.’ I dunno. ChatGPT Health. GPT-5.5-Instant optimizes for health advice. Middle Of The Journey. More on the MidJourney scanner. New Medical Diagnostic Just Dropped. Another new AI diagnostic system. Google on AI Control. Laying out the basics of defense in depth. The Once And Future Fable. Planning for a measure of severity. Fable: The First Lawsuit. A customer sues for access, has some good points. Dean Ball Joins OpenAI. This is quite the upgrade all around. Show Me the Money. Quiet week. Taste Labs raises $18.5m. Quiet Speculations. Bets about future compute prices. Alex Bores Loses In NY-12 By 4%. We almost got there. The Quest for Sane Regulations. The doors are now open. Chip City. An ASML machine is missing. The Week in Audio. Donald Trump on Anthropic, Ball on Labenz, Clark Odd Lots. People Just Say Things. Rhetorical Innovation. Know the rules. There Are Two Pills. Are you only AGI pilled? Or are you ASI pilled? Who Evals The Evals. The quest for eval consensus. Aligning a Smarter Than Human Intelligence is Difficult. Beneficial traits. Cooperative Alignment. Opus 4.7 and 4.8 are not distillations. People Are Worried About AI Killing Everyone. Francis Fukuyama. Other People Are Not As Worried About AI Killing Everyone. Alas, DC. The Lighter Side. Not cool, man. Language Models Offer Mundane Utility Automatically update and fix old academic papers, such as your own. Help clinicians revisit unsolved rare pediatric disease cases, and that’s with o3. Tokens are cheap, so if you can loop over useful things, you do it. /goal /loop. Tom Osman: This “loop” automation is nuts inside of Codex. “/goal go over every single feature in this app create a user story with expected behaviour based on the code keep a single canonical spreadsheet tracking the features status – when done switch loop to testing every user story and documenting all errors – when done fix every logistical error or ux error – test every user behaviour again post fix” Shoutout to @MatthewBerman for the heads up. Hundreds of user stories being worked through like it’s nothing. Use Mercury’s new Command feature to set details for a wire? Definitely scary stuff the first times you do it purely for error reasons, and after that also for potential prompt injection reasons. The human check step will stick around for a while. Keep your company or other group small. Paul Graham: One of the biggest advantages of AI will be that it lets companies get further before they cross the lines (at about 10 and about 150 people) beyond which groups become less productive. This leaves out the biggest thresholds to avoid, which are 2 and 3. Grok, like the internet, is for porn. Two former engineers estimated that adult content was the majority of usage. Language Models Don’t Offer Mundane Utility European parliament, to ‘reduce dependence on American technology firms,’ scraps Google search for the French Qwant, which is still substantially dependent on Bing. Huh, Upgrades Claude Code now supports Artifacts, starting with Team and Enterprise plans. We have another new version of GPT-5.5-Instant. With all due respect and thanks for what I presume are small improvements, if you are changing the model change the version number, why is this hard, v5.5.1 is a thing you can do. On Your Marks Refine ‘win 90% of the time head-to-head against AI reviewers’ on economics preprints (and those in other related areas), including beating Fable. There was no human comparison. Ben Golub: A system won if it identified more genuine concerns the other system missed (“residual concerns”). Refine averaged 28.1 unique residual concerns per match; comparison reviews averaged 14.5. Substantive concerns: 22.1 vs. 11.8. There are ways for Refine to ace that benchmark without it meaning much, although I have also heard good things from other tests, including from Tyler Cowen. I presume it is indeed good at finding potential issues in economics papers, but I also presume it would not be that much effort to make Fable similarly good at this. Deepfaketown and Botpocalypse Soon Undersecretary of State Jacob Helberg has AI write his article about ‘The Digital Sovereignty Trap’ warning other countries not to build their own sovereign AI systems. I confirmed this with Pangram, and this finally convinced me to stop holding out and sign up for Pangram, but also I very much did not need Pangram. Also most of this: Teortaxes: This is pretty much a supervillain’s speech The only way he could twist the knife more is if he just said Europeans are natural slaves (like I do). Why “apply the heresy to nations”? The whole point of the EU is a shared economic and policy bloc that can compete. “No, you can’t.” Andrew Curran: It does sound like a villain’s monologue scene honestly. It has the rhythm. For now Helberg got the main concrete thing he wanted here: Europe signed on to Pax Silica, which is a program to ensure countries integrate with the American AI supply chain and not the Chinese one. By contrast, Rubio’s Views on America is fully human written. He does it all. To be fair to all sides, here’s Hunter Biden using AI to write his reaction to the NY elections, again we did not need Pangram, are you kidding me. Again, we learn not that AI is a good writer, or that humans are bad writers, but that the literary prize judgment processes are worthless. Jack: That which can be won with undisclosed AI output should be Nabeel S. Qureshi: *Another* apparently AI-generated story wins a literary prize, this time judged by a panel including the novelist Ruth Ozeki. Literary prizes need to start including Pangram checks in their process, or else change the rules to make AI writing ok. It’s very simple! Nat McAleese: the fact that slop has won a couple of literary prizes implies that slop is in large part an exposure effect; more broadly frequent AI users probably underestimate how good slop is by 2019’s standards. (story is ass imo) In theory sure you can imagine AI writing being good if you lacked the taste to recognize that it is slop. But almost always no, also I can’t even with this one. This is how it starts: Back and Forth By Kavyta Kay The tree knew before she did – and it waited. Deepa would realise this later, when clarity r
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力