跳到主内容
@wquguru
精选85r/LocalLLaMA(Reddit)模型发布/更新

UkisAI发布Swift系列27B模型,推理Token减少63%

UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy

原文
发到 X
推荐理由

专注优化推理效率的本地模型新选择,大幅降低思考Token同时保持高精度,适合对延迟敏感的场景。

Hey everyone,

大家好,

Jovan from UkisAI here! Today, we are introducing Swift, a family of efficient reasoning LLMs based on Qwen, trained by penalizing tokens related to pathological overthinking patterns and restoring accuracy via RL (GSPO) and OPD.

我是来自 UkisAI 的 Jovan!今天,我们介绍 Swift,这是一系列基于 Qwen 的高效推理大语言模型(LLM),通过惩罚与病态过度思考模式相关的 token,并利用强化学习(GSPO)和 OPD 恢复准确性进行训练。

After amazing feedback and 350k+ downloads in 13 days on our Swift Qwen 3.8 27B we are releasing the entire model family as well as the highly requested GSQ-RCO quants for 27B and Flash-Next.

在 Swift Qwen 3.8 27B 于 13 天内获得惊艳反馈和超过 35 万次下载后,我们现在发布整个模型家族以及备受期待的 GSQ-RCO 量化版本(针对 27B 和 Flash-Next)。

This release includes:

本次发布包括:

Swift1.5 27B, an improved version of our last model, with even lower token usage, fixed bugs and better agentic performance, with -58.5% thinking tokens while scoring 0.35% higher and outperfoming base on Terminal Bench 2.1 by not falling into "overthinking error" loops.

Swift1.5 27B,这是我们上一款模型的改进版,具有更低的 token 使用量、修复了 bug 并提升了智能体性能,思考 token 减少了 58.5%,同时得分高出 0.35%,并且在 Terminal Bench 2.1 上表现优于基础模型,因为它不会陷入“过度思考错误”循环。

Swift Flash Next, with 63.4% fewer thinking tokens and a 1.8x speed up scoring -0.2% vs base on xhigh

Swift Flash Next,思考 token 减少了 63.4%,速度提升 1.8 倍,在 xhigh 基准测试中得分比基础模型低 0.2%。

Swift Bonsai 2, with 39.8% fewer thinking tokens while scoring 0.19% higher (although we'd still like to note it as experimental)

Swift Bonsai 2,思考 token 减少了 39.8%,同时得分高出 0.19%(尽管我们仍希望将其标注为实验性版本)。

Our benchmarks are ran x5 on Base and Swift, averaging across five seeds and various domains, including General (GPQA, AIME26), Coding (LiveCodeBench), Vision (ERQA), Agentic (Terminal Bench 2.1).

我们的基准测试在 Base 和 Swift 上运行了 5 次,对五个随机种子和各种领域进行了平均,包括通用领域(GPQA, AIME26)、编程(LiveCodeBench)、视觉(ERQA)和智能体(Terminal Bench 2.1)。

One note is that the Terminal Bench 2.1 score of Swift1.5 27B is misleadingly low at first glance. It is not a bug, but a simple matter of the Swift models not falling into overthinking loops and failing the task, rather pursuing it until the end, leading to higher average token usage. The token reduction still falls in the -38.7% range when compared apples-to-apples.

需要注意的是,Swift1.5 27B 在 Terminal Bench 2.1 上的得分乍一看似乎偏低。这不是 bug,而是因为 Swift 模型没有陷入过度思考循环导致任务失败,而是坚持完成任务直到结束,从而导致平均 token 使用量更高。当进行公平比较时,token 减少幅度仍在 -38.7% 左右。

We also added a fun "game creation" benchmark you can find and play here, it is completely subjective but Swift generated better games in less time: Flash Next Game and 27B Game

我们还添加了一个有趣的“游戏创建”基准测试,你可以在这里找到并试玩,虽然这完全是主观的,但 Swift 用更少的时间生成了更好的游戏:Flash Next Game 和 27B Game。

We are including a Research API and HuggingFace Spaces to give the models a spin before downloading or if you don't have enough compute to run them right now! You can find both on the model cards.

我们提供了 Research API 和 HuggingFace Spaces,以便你在下载之前或目前计算资源不足以运行模型时体验这些模型!你可以在模型卡片中找到它们。

We have also made GGUF, NVFP4, MLX and W4A16 quants for relevant model versions.

我们还为相关模型版本提供了 GGUF、NVFP4、MLX 和 W4A16 量化版本。

More details on our training approach and community feedback can be seen here: https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai_swiftqwen3827b_583_thinking_x195_speed/

更多关于我们的训练方法和社区反馈的详情见此处:https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai_swiftqwen3827b_583_thinking_x195_speed/

All of the various quantization and model versions are available in their respective collections:

各种量化和模型版本可在各自的集合中找到:

Swift1.5 27B: https://huggingface.co/collections/ukisai/swift-15-27b

Swift Flash Next: https://huggingface.co/collections/ukisai/swift-flash-next

Swift Bonsai 2: https://huggingface.co/collections/ukisai/swift-bonsai-2

We are also working on a 9B variant to be released in the upcoming days.

我们还在开发一个9B变体,将在未来几天内发布。

We would greatly appreciate your feedback via independent evaluations on real world tasks. As per last release, we operate on a candy-shop basis, trying to fulfill as many Swift model requests and quants as possible, so please do share your needs in the comments!

我们非常感谢您通过独立评估对真实世界任务提供的反馈。根据上次发布的做法,我们以糖果店模式运作,尽力满足尽可能多的Swift模型请求和量化需求,因此请在评论中分享您的需求!

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件