跳到主内容
@wquguru
精选75Two Minute Papers(YouTube)模型发布/更新多源精选 ×9

Qwen 3.8 Flash Next 评测:开源模型挑战闭源巨头

This Free AI Just Caught The Billion Dollar Giants

原文
发到 X

This is the new Quen 3.8 Flash Next AI and the first results look fantastic. Incredible system and we are lucky to get so many of these open and free AI systems that it's easy to confuse which is which fellow scholars that is a very good problem to have and it seems that we can switch out paid closed systems to open weights AI that we can download and run ourselves forever. no subscriptions and we have a paper too. We are spoiled here and I am really grateful and that is why I keep covering as many of these as I can.

I want to show you how I and you fellow scholars are already using it out there in the wild. Now there is Quen 3.8 Max incredibly huge. Then the 27 billion. This runs on a very beefy laptop or a desktop. This is 3.8 flash next. This is in the middle but it is so different. The small guy 27B is a dense model. Now this one is a mixture of experts model. Very different. This means that for each token it only uses a small part of the model.

Great for systems with lots of memory and slower memory bandwidth like the DJX Spark or two. I am running experiments with it on two sparks. Day one. It runs about 38 tokens per second. Very respectable. Thank you so much to Jensen, Brad, and Mark from Nvidia for the hardware. Now, you brilliant fellow scholars already found a way to run this on much more modest hardware. Now, 3.8 flash next is the first of their next generation model and has substantial innovations.

Here's three. Dear fellow scholars, this is two minute papers with Dr. Koa Eher. One, when sparse attention as we start adding documents or just keep talking to the AI longer, its memory we call context starts filling up and with that things get really expensive. This is called full attention and it has quadratic complexity. So processing twice as much context is roughly four times as much work. Deepseek tried to ease this with DSA which picks only important individual tokens.

Quen's new QSA however goes further. It bundles these tokens up into tiny blocks and only searches those blocks. Why? Well, because it makes this kind of growing context even cheaper. Two, for each token, the model carries additional information about that token. But here the layers can step on each other's toes because they all keep rewriting the same running information. Now here instead they use not one but four branches.

Why? Because here information can be left alone while the rest is changed. They call it gated residual. And three imagine a hot dog. Yummy. Now the two words separately mean something quite different, right? Hot and dog does not equal hot dog. So Quen can bunch these short token combinations together and build lookup memory for them to retrieve from quickly and cheaply. They call it engram embedding. It's a variant of what DeepS already does.

Deepseek spreads this lookup memory across multiple layers while Quen puts it into one huge lookup layer near the start. And these three give you a system that already outperforms some of the best open weights AI systems, maybe even DeepSeek 4 Pro that is much much bigger. And all this just a few days after release. This is mindblowing. And we get all this for free forever. What a time to be alive. We need new tools for the era of LLMs.

And Weights and Biases now has Weave, a lightweight toolkit to confidently iterate on LLM applications. Use traces to debug how data flows through each step of your app and use evaluations to measure your progress. It is the best. Try it out now at wnb.me/papers me/papers or click the link in the description below.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件