跳到主内容
@wquguru
精选80Rohan Paul模型发布/更新多源精选 ×3

NVIDIA 发布 Nemotron-3 嵌入模型系列,8B 模型 RTEB 排名第一

A new family of open, commercially available embedding models built for agentic…

原文
发到 X

A new family of open, commercially available embedding models built for agentic retrieval, code retrieval, and agent memory.

NVIDIA Nemotron-3-Embed-8B-BF16 now ranks first on RTEB with a 78.5% retrieval score.

RTEB is a RAG retrieval leaderboard, showing how reliably a model finds the right context.

When retrieval fails, agents get irrelevant evidence, leading to more searches, longer reasoning and higher token bills.

NVIDIA Nemotron 3 Embed breaks that bottleneck with one 8B model and two 1B alternatives.

The NVIDIA Nemotron-3-Embed-1B-BF16 achieved a 72.4% score on RTEB, a 27% reduction in error rate over its predecessor.

A full family across the accuracy-efficiency curve - the 8B flagship plus two efficient 1B variants: 1B-BF16 for cost- and latency-sensitive serving, and 1B-NVFP4, a Blackwell-optimized build with up to 2x higher throughput that retains 99%+ of BF16 accuracy.

The 1B version, Nemotron-3-Embed-1B-NVFP4 maintains over 99% BF16 accuracy and can handle up to twice as much BF16 traffic.

All three models can take up to 32K tokens as input, which can be long documents, code or agent histories.

Created by NVIDIA using structured pruning, teacher distillation and longer and longer training contexts, NVIDIA Nemotron-3-Embed-1B-BF16

Tests with the NVIDIA Nemotron 3 Ultra showed that improved retrieval led to fewer searches and lower token costs.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →