跳到主内容
@wquguru
精选92OpenAI(YouTube)模型发布/更新

OpenAI发布GPT-6 Astra,用户实测展现高可靠性与自然交互

GPT-6 Astra with Peter Gostev

原文
发到 X
推荐理由

OpenAI正式推出代号为Astra的新模型,开发者实测显示其在代码生成可靠性和交互自然度上有质的飞跃,这是值得关注的重大产品迭代。

I value this so much that I don't need to babysit it because I only have so much space in my head and I cannot think of every single task. How did you think about what to evaluate Astra on? If I give it like a medium difficulty test, any model can do it or so many models can do it now. So I have to think of what's the hardest thing I can push it on and make it work for as long as possible and really explore the space.

我非常重视这一点,因此我不需要像保姆一样盯着它,因为我的大脑容量有限,无法记住每一个任务。你是如何思考评估 Astra 的标准?如果我给它一个中等难度的测试,任何模型都能做到,或者说现在有很多模型都能做到。所以我必须思考我能让它应对的最困难的任务是什么,并尽可能长时间地让它工作,真正探索其能力边界。

Yeah, so what I've been trying to do, and I really like this kind of 3D visual tasks, but I like them because I can visually track how capabilities of the models improve. And now what I've built here is a voxel 3D representation of historic London, which goes across the different eras, medieval London, Tudor London, and it transforms all within the same map. But here I pretty much gave it maybe a couple more prompts in terms of I wanted to go overhead so it looks like GTA 2.

是的,我一直在尝试做的事情是,我非常喜欢这类 3D 视觉任务,我喜欢它们是因为我可以直观地追踪模型能力的提升。而我现在在这里构建的是伦敦历史的体素 3D 表示,跨越了不同的时代:中世纪伦敦、都铎王朝伦敦,并在同一张地图中进行转换。但在这里,我基本上只给了它几个额外的提示,比如我希望从俯视角度查看,使其看起来像 GTA 2。

I would say with Sol, one thing that felt a little bit off is that anytime you give it some kind of criticism or feedback or something like that, it would instantly say, oh yes, that's great. I actually thought exactly the same thing. You're so right. In here, I would say with Astra, it really felt kind of a shift. I didn't feel anything annoying about the communication, but it felt much more natural, like a human would push back or accept that it made a mistake.

我会说在使用 Sol 时,有一点感觉不太对劲:无论何时你给它某种批评或反馈之类的东西,它都会立刻说:“哦,是的,太棒了。我其实也完全这么想。你说得太对了。”而在 Astra 这里,我感觉到了明显的转变。我觉得它的沟通方式没有任何令人讨厌的地方,而是更加自然,就像人类会反驳或承认自己犯错一样。

So the one problem I keep trying to solve is I've got a little app that I've vibe-coded probably since the days of, I know, GPT 5.2 kind of time. So the code is kind of bad and it's accumulating over time. And I think I've got probably like 150,000 lines of code in it. And I keep trying to migrate it or improve it. With 5.6, it worked, but I had to do a lot of debugging later to align it. But with Astra, it pretty much just worked and I didn't need to do anything.

所以我一直在试图解决的一个问题是:我有一个小应用,可能是从 GPT 5.2 那个时期开始用 vibe coding(氛围编码)写成的。所以代码质量不太好,而且随着时间推移不断累积。我想里面大概有 15 万行代码。我一直试图迁移或改进它。使用 5.6 版本时虽然能工作,但我后来做了大量调试才能对齐。但使用 Astra 时,它几乎直接就能工作,我不需要做任何事情。

So for me, that jump in reliability that I don't need to babysit it with every single step, that's a big deal. So what are you bound by now? So CPU is surprisingly the big one. I had to move things off of my laptop to a Linux box. So it works remotely, it's plugged in, battery doesn't die. So that's the big deal now. Yeah, it really feels like the model's grown up.

所以对我来说,这种可靠性的飞跃意味着我不需要在每一步都像保姆一样盯着它,这是一件大事。那么你现在受限于什么? surprisingly CPU 成了主要瓶颈。我不得不把一些东西从笔记本电脑移到一台 Linux 主机上。所以现在它是远程运行的,插着电源,电池不会耗尽。这就是现在的关键所在。是的,真的感觉这个模型已经成熟了。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近