跳到主内容
@wquguru
精选75AI Explained(YouTube)产品发布/更新

Anthropic发布Claude Co-work,AI自动化白领工作引争议

Anthropic: Our AI just created a tool that can ‘automate all white collar work’, Me:

原文
发到 X

the CEO of one of the major AI labs predicted last year that by around now 100% of the code written by that company would be produced by one of their AI models. Next up within 2026 would be all other knowledge work and a new tool released by Anthropic in the last couple of days seems to back that up. It's called Claude Co-work. Not only has it gone omega viral at 42 million views for its ability to automate non-coding tasks, the tool itself was produced within Clawude code powered by their latest frontier model, Claude Opus 4.5, thereby seeming to justify the prediction that essentially all of the code would by now be written by AI.

So wait, if they got that right, does that mean that Anthropic and those like Schultto Douglas are correct when they say that in 2026, this year the same will be true of automating all white collar work.

The most striking thing about next year is that the other forms of knowledge work are going to experience what software engineers are feeling right now where they went from typing, you know, most of their lines of code at the beginning of the year to typing barely any of them at the end of the year. I think of this as the Claude code experience, but for all forms of knowledge work. I also think that probably continual learning gets solved in a satisfying way.

Well, I've been using Claude code for quite a while and yes, have been playing about with the new Claude co-work. And for me, those predictions are just not true. But so many of us might then throw the baby out with the bath water and miss out on some pretty crazy productivity gains. So, I'm going to show why we shouldn't underestimate the gains to be had either. Then, for those who want to go a bit deeper, I'm going to end with the why.

Why can models produce genius like seeing tiny bugs in large code bases and writing for me powerful poems but also still fail at such basic tasks? No, I don't mean how many A's in the word orange. Although surprisingly GPT 5.2 still can't get that right. No, I mean why are they still sometimes so brittle memorizing that Tom Smith's wife is Mary Stone but not deducing that Mary Stone's husband is Tom Smith? And what does any of this mean for your job, white collar or otherwise?

What does the latest data show? First of course, a quick word on Claude Co-work, which inevitably it seems some people are calling AGI. This, of course, follows numerous viral posts and articles about the underlying model Claude Opus 4.5 when given the right scaffold already being AGI. Indeed, a long list of notable commentators have this perspective. These posts can lead, of course, to two very desperate reactions, both of which I'd advise against.

One that it's all BS, all hype merchants. these tools hallucinate all the time and are pretty much useless and second that they are AGI perhaps and I you are just missing out. We can't understand how to use them. We're missing out so much our careers are doomed. This video is hopefully going to channel you down the middle path which is you can get great productivity gains but they're not there yet. For context, I've been using Claude Code for a very long time and co-work for the last 48 hours. to slightly debunk the hype point.

If I gave a new employee this task, make a comparison chart for this football club's league position at this date today for each of the last five seasons. Add it as a PowerPoint to my desktop. Oh, and ask any clarifying questions and share a plan of how you'll approach this task. I would expect, and let me know if you agree or not, for them to either say at the end of the day, I couldn't find any source to give definitive answers on that question or to have produced the relevant PowerPoint.

Now, you can see the co-work tab here and the kind of questions it lays out, and it does indeed give a great plan. I approved it immediately, and it didn't even take that long to be honest. The result, I would say, was visually quite impressive and pretty much acceptable. Obviously, you have to pick a moderately hard task because if it's too easy, you just do it yourself. So, this was the result. Slight problem. I checked two of the dates it gave me for January 2023 and 2025 and the league position of this club, and both were incorrect.

I manually checked and within about 5 minutes I found two other data sources BBC and this site 11v11 both of which said that Stockport were seventh at the time not third for January 13th 2025. This co-working AGI by the way did not caveat its results in its summary to me that it couldn't find a reliable source either. Now, I could of course give you hundreds of such examples from the legendary Claude code powered by Claude Opus 4.5, but that wouldn't be too interesting or fair on you cuz you'd have to see the whole context of the codebase.

I just don't want you guys to walk away from these viral posts thinking unless I spend all my money and keep up with a tool released just last week, I'm going to completely fail at my white collar job. And if the models make any mistakes, I'm the dumb one. I must have done something wrong. But I don't want you to make the opposite mistake, which is to completely ignore these tools and think that they can't boost your productivity at all.

The truth lies somewhere in the middle. And look, even the lead developer for Claude Code said as much later on in a reply after saying all of the code for Claude Co-work was written by Claude Opus 4.5. He clarified, "It was not zero intervention. We the humans had to plan, design, and go back and forth with Claude. Which then for my super smart audience leads to a key question. Well, is it faster to get Claude code to do the draft and then reddraft and then test fail reddraft and then kind of get it right or for the human to just do it themselves from scratch, whether it be coding or just other white collar work.

Thankfully, we have a key clue from this OpenAI paper from October of 2025. Using blind human grading, we have already passed that tipping point. We get more of a productivity multiplier by getting models to try again and again and again and the human to just step in, review, and edit than from the human just doing it themselves. This GDP valve paper covers dozens of white collar industries and I did an entire video on it.

So I'm not going to go into too much depth, but that for me is the real tipping point. And yes, I've experienced that in my own coding, which I do almost every day. It makes a bunch of dumb and sometimes dangerous mistakes, but don't throw the baby out with the bath water. Even take my stopport PowerPoint. It's really quite welldesigned and almost all the other facts are true. So I could just edit a couple of the numbers and have a decent presentation in less time than creating it myself from scratch.

Quick bit of technical detail. Claude co-work is only available on the max tier. Minimum $90 or $100 and on max only. That's Mac OS, not Macs. Mac OS, not Windows. But also Max only, not the pro tier of Claude. Notice this productivity speed up though is only true for a certain number of the latest models. those most likely to be tried by enthusiasts like us, less so the general population, and also that those models like GBTC 5.2 Pro or 4.5 Opus are also gated heavily by price.

If we are right about that tipping point and about how few people are using the latest models with the best scaffolds, then you'd expect the current AI impact on productivity and the labor market to be relatively limited. And what does the data show according to this January 7th, 2026 report from the widely cited Oxford Economics? Well, to me, it shows exactly that. Yes, new graduates face slightly higher unemployment, but that isn't out of line with other historical trends.

If you're listening to this, the new graduate unemployment rate has been much higher in the quite recent past, like 2015 or 2010. The authors note, if you zoom in on this graph, that there's actually been a slight downward tre

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近