跳到主内容
精选85Rohan Paul模型发布/更新

295B模型压缩至1位,本地运行比云API快2.2倍

This is so interesting for local LLM.

原文
推荐理由

本地大模型玩家注意了,1-bit 压缩让 295B 模型跑在消费级显卡上,速度反超云端,值得复现测试。

This is so interesting for local LLM.

Atomic Chat compressed the 295B-parameter Hy3 (to an extreme 1-bit ) model into a 92GB file and ran it 2.2x faster than its cloud API.

They installed the model locally on 4x RTX 5090 with 128GB VRAM against the same Hy3 over cloud API.

The big deal is that the extreme compression kept useful quality after this extreme compression.

Both versions built working Flappy Bird, Arkanoid, and Snake games from one-shot prompts, with no crashes or obvious visual gap.

Outputs: Hy3 1-bit local: 76.9K tokens, 15.5 min Hy3 cloud API: 75.1K tokens, 34.3 min

Hy3 is Tencent's open-sourced Apache 2.0 licensed model, i.e. commercially usable by enterprises. A massive 295B MoE (21B active parameters).

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近