跳到主内容
@wquguru
精选70The Decoder(RSS)模型发布/更新

谷歌DiffusionGemma:用不到10%预算将Gemma改造成扩散模型

Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model

原文
发到 X

Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second. Quality still trails the original autoregressive model in benchmarks, especially on reasoning tasks.

谷歌DeepMind没有从头训练新模型,而是用不到原始训练预算10%的成本,将Gemma 4改造为扩散模型。DiffusionGemma并行生成256个令牌,而非逐个生成,每秒约处理1500个令牌。在基准测试中,其质量仍不及原始自回归模型,尤其在推理任务上。

The article Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model appeared first on The Decoder.

文章《谷歌的DiffusionGemma证明构建文本扩散模型无需从头训练》首次出现在The Decoder上。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近