Google DeepMind发布EmbeddingGemma 2多模态嵌入模型
Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4
面向端侧的多模态嵌入模型落地细节详实,量化内存占用极低且支持多种模态混合检索,做RAG或本地搜索的同学值得重点关注。
Google DeepMind has released EmbeddingGemma 2, an open model that embeds text, code, images, video and audio into one 768-dimensional space. It has 740M parameters, an 8K token context window and an Apache 2.0 license. It targets on-device search, classification and privacy-first RAG. This article analyzes, compares and showcase how EmbeddingGemma 2 fits in the space.
Google DeepMind 发布了 EmbeddingGemma 2,这是一个开源模型,能将文本、代码、图像、视频和音频嵌入到一个 768 维的空间中。它拥有 7.4 亿参数、8K token 上下文窗口,并采用 Apache 2.0 许可证。它面向设备端搜索、分类以及隐私优先的 RAG(检索增强生成)。本文分析、比较并展示 EmbeddingGemma 2 在该领域的应用情况。
Deployable today? Yes. Weights are live on Hugging Face and Kaggle, with Ollama, llama.cpp GGUF and LiteRT builds available now.
今天可以部署吗?可以。权重已上线 Hugging Face 和 Kaggle,目前提供 Ollama、llama.cpp GGUF 和 LiteRT 构建版本。
What an Embedding Model Does
嵌入模型的作用
An embedding model converts content into a vector of numbers that captures meaning. Similar items land close together, so they are easy to search and compare. In a RAG pipeline, these vectors let an LLM retrieve fresh information it was not trained on. Generating embeddings locally keeps data on the device, cuts latency and works offline.
嵌入模型将内容转换为捕捉含义的数字向量。相似的项目在向量空间中彼此靠近,因此易于搜索和比较。在 RAG 管道中,这些向量使大语言模型(LLM)能够检索其训练数据之外的新鲜信息。在本地生成嵌入可将数据保留在设备上,降低延迟,并支持离线工作。
One Vector Space for Every Modality
所有模态的统一向量空间
EmbeddingGemma 2 is built on the Gemma 4 architecture. A text query can retrieve a photo. A voice memo can retrieve a video clip. Interleaved inputs, like a product listing with text, images and a demo video, produce a single embedding.
EmbeddingGemma 2 基于 Gemma 4 架构构建。一条文本查询可以检索照片。一段语音备忘录可以检索视频片段。混合输入(如包含文本、图像和演示视频的产品列表)会产生单个嵌入向量。
The design is modular. It has three parts:
该设计是模块化的。它由三个部分组成:
- Text and code backbone: 270M parameters (130M transformer plus 140M embedder)
- Vision encoder: 170M parameters, optional
- Audio encoder: 300M parameters, optional
- 文本和代码主干:2.7 亿参数(1.3 亿 Transformer 加 1.4 亿嵌入器)
- 视觉编码器:1.7 亿参数,可选
- 音频编码器:3 亿参数,可选
Developers load only what they need: 270M for text, 440M for text and vision, 570M for text and audio, or 740M for everything. All setups share one vector space. A query embedded with the text-only setup can match documents embedded by the full model.
开发者只需加载所需部分:文本为 2.7 亿,文本加视觉为 4.4 亿,文本加音频为 5.7 亿,全部功能为 7.4 亿。所有配置共享同一个向量空间。使用仅文本配置嵌入的查询,可以与由完整模型嵌入的文档进行匹配。
The context window is 8,192 tokens, 4x larger than version 1. That fits about 29 images, 58 video frames or 5.5 minutes of audio.
上下文窗口为 8,192 个 token,比版本 1 大了 4 倍。这大约能容纳 29 张图像、58 帧视频或 5.5 分钟的音频。
Benchmarks
基准测试
Google research team reports leading scores among sub-1B multimodal embedders on MTEB Code and MAEB. Full-precision results at 768 dimensions:
Google 研究团队报告称,在 MTEB Code 和 MAEB 上,其得分在参数量小于 10 亿的 multimodal embedders 中领先。768 维全精度结果如下:
| Benchmark | EmbeddingGemma 2 | EmbeddingGemma 1 |
|---|---|---|
| MTEB multilingual v2 | 61.36 | 61.15 |
| MTEB Code v1 | 78.68 | 68.76 |
| MIEB lite (image) | 64.64 | n/a |
| MMEB v2 overall | 59.01 | n/a |
| MSEB retrieval (sound) | 69.54 | n/a |
| MAEB (audio) | 49.39 | n/a |
| 基准测试 | EmbeddingGemma 2 | EmbeddingGemma 1 |
|---|---|---|
| MTEB multilingual v2 | 61.36 | 61.15 |
| MTEB Code v1 | 78.68 | 68.76 |
| MIEB lite (image) | 64.64 | n/a |
| MMEB v2 overall | 59.01 | n/a |
| MSEB retrieval (sound) | 69.54 | n/a |
| MAEB (audio) | 49.39 | n/a |
Source: EmbeddingGemma 2 model card
来源:EmbeddingGemma 2 模型卡片
Code retrieval gains 9.92 points, roughly 14%. Multilingual text quality holds steady. Bigger models still lead some boards. Qwen3-VL-Embedding-2B reports 73.2 on its own MMEB-V2 run, with about 2.7x the parameters and no audio support.
代码检索得分提升 9.92 分,约为 14%。多语言文本质量保持稳定。更大的模型在某些榜单上仍占优势。Qwen3-VL-Embedding-2B 在其自身的 MMEB-V2 运行中报告得分为 73.2,参数量约为前者的 2.7 倍,且不支持音频。
Built for Phones and Laptops
专为手机和笔记本电脑打造
With quantization on a Pixel 11 Pro, active RAM is about 191MB for text-only weights. The full multimodal model needs about 567MB. Quantization-aware training compresses weights to INT4 and INT8. The Google AI Edge team measured 37.3 ms per image on a MacBook M5 Pro GPU, using a 70-token vision budget.
在 Pixel 11 Pro 上进行量化,纯文本权重的活跃 RAM 约为 191MB。完整的多模态模型需要约 567MB。感知量化的训练将权重压缩为 INT4 和 INT8。Google AI Edge 团队在配备 M5 Pro GPU 的 MacBook 上,使用 70 个 token 的视觉预算,测得每张图像的处理时间为 37.3 毫秒。
Matryoshka Representation Learning (MRL) lets developers truncate vectors to 512, 256 or 128 dimensions. Moving from 768 to 128 dimensions cuts storage up to 6x. At 256 dimensions, MTEB multilingual only slips from 61.36 to 60.41. At 128 dimensions, MMEB drops to 45.65, so Google recommends 128d mainly for text-only workloads.
Matryoshka Representation Learning (MRL) 允许开发者将向量截断为 512、256 或 128 维。从 768 维降至 128 维可减少高达 6 倍的存储空间。在 256 维时,MTEB 多语言性能仅从 61.36 微降至 60.41。在 128 维时,MMEB 降至 45.65,因此 Google 建议 128d 主要用于纯文本工作负载。
Interactive Explainer
交互式解释器
EmbeddingGemma 2 vs. Closest Competitors
EmbeddingGemma 2 与最接近的竞争者对比
| Feature | EmbeddingGemma 2 | EmbeddingGemma 1 | Qwen3-VL-Embedding-2B | LCO-Embedding-Omni-3B | Gemini Embedding 2 |
|---|---|---|---|---|---|
| Developer | Google DeepMind | Google DeepMind | Alibaba Qwen | LCO-Embedding (research) | |
| Parameters | 740M (270M text-only) | 308M | 2B | 3B backbone (5B listed on HF) | Not disclosed |
| Text / code | Yes | Yes | Yes | Yes | Yes |
| Images | Yes | No | Yes | Yes | Yes |
| Video | Yes | No | Yes | Yes | Yes |
| Audio | Yes | No | No | Yes | Yes |
| Output dims (MRL) | 768 (512, 256, 128) | 768 (down to 128) | Up to 2048 (64 to 2048) | Not stated | 3072 (128 to 3072) |
| Context | 8,192 tokens | 2K tokens | 32K tokens | Not stated | 8,192 tokens |
| Languages | 100+ | 100+ | 30+ | Not stated | 100+ |
| License / access | Apache 2.0, open weights | Open weights (Gemma terms) | Apache 2.0, open weights | Apache 2.0, open weights | Paid API only |
| Published on-device RAM | ~191MB text, ~567MB full | Under 200MB | Not published | Not published | Cloud only |
| Source | Model card | Docs | HF card | HF card | API docs |
| 特性 | EmbeddingGemma 2 | EmbeddingGemma 1 | Qwen3-VL-Embedding-2B | LCO-Embedding-Omni-3B | Gemini Embedding 2 |
|---|---|---|---|---|---|
| 开发者 | Google DeepMind | Google DeepMind | Alibaba Qwen | LCO-Embedding (研究) | |
| 参数量 | 740M (270M 纯文本) | 308M | 2B | 3B 主干 (HF 上列出为 5B) | 未公开 |
| 文本 / 代码 | 是 | 是 | 是 | 是 | 是 |
| 图像 | 是 | 否 | 是 | 是 | 是 |
| 视频 | 是 | 否 | 是 | 是 | 是 |
| 音频 | 是 | 否 | 否 | 是 | 是 |
| 输出维度 (MRL) | 768 (512, 256, 128) | 768 (最低至 128) | 最高 2048 (64 至 2048) | 未说明 | 3072 (128 至 3072) |
| 上下文 | 8,192 tokens | 2K tokens | 32K tokens | 未说明 | 8,192 tokens |
| 语言 | 100+ | 100+ | 30+ | 未说明 | 100+ |
| 许可 / 访问方式 | Apache 2.0, 开放权重 | 开放权重 (Gemma 条款) | Apache 2.0, 开放权重 | Apache 2.0, 开放权重 | 仅限付费 API |
| 已发布的端侧 RAM | ~191MB 文本, ~567MB 完整 | 低于 200MB | 未发布 | 未发布 | 仅限云端 |
| 来源 | 模型卡片 | 文档 | HF 卡片 | HF 卡片 | API 文档 |
How to Run It
如何运行
It runs on sentence-transformers v6.1.0+, Transformers, vLLM, SGLang, MLX, llama.cpp, Ollama, LM Studio, LiteRT and MediaPipe. Qdrant covers vector storage and Unsloth covers fine-tuning. ML Kit support for Android, with NPU acceleration, is coming within weeks.
它支持 sentence-transformers v6.1.0+、Transformers、vLLM、SGLang、MLX、llama.cpp、Ollama、LM Studio、LiteRT 和 MediaPipe。Qdrant 用于向量存储,Unsloth 用于微调。Android 的 ML Kit 支持(含 NPU 加速)将在几周内推出。
pip install -U "sentence-transformers[image,audio,video]" transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2")
q = model.encode("What causes the northern lights?", prompt_name="SearchQuery")
d = model.encode("Charged particles from the sun.", prompt_name="Document")
print(model.similarity(q, d))On Ollama, run ollama pull embeddinggemma-2. Tags range from 270m (378MB) to 740m (1.3GB). Demos live in Google AI Edge Gallery. See the developer guide for more.
在 Ollama 上,运行 ollama pull embeddinggemma-2。标签范围从 270m (378MB) 到 740m (1.3GB)。演示位于 Google AI Edge Gallery。更多详情请参阅开发者指南。
Key Takeaways
关键要点
- One 740M open model embeds text, code, images, video and audio into a shared 768d space.
- Modular encoders scale the footprint from 270M (text) to 740M (full multimodal).
- Code retrieval jumps from 68.76 to 78.68 on MTEB Code.
- Runs in ~191MB to ~567MB of RAM on a Pixel 11 Pro with quantization.
- 一个 740M 参数的开源模型可将文本、代码、图像、视频和音频嵌入到共享的 768d 空间中。
- 模块化编码器将模型规模从 270M (文本) 扩展至 740M (完整多模态)。
- 在 MTEB Code 上,代码检索性能从 68.76 跃升至 78.68。
- 在 Pixel 11 Pro 上使用量化时,RAM 占用约为 ~191MB 至 ~567MB。
FAQ
常见问题解答
- Can EmbeddingGemma 2 be used commercially? Yes. It is released under the Apache 2.0 license.
- How much memory does EmbeddingGemma 2 need? Google reports about 191MB of active RAM for text-only use and 567MB for full multimodal use, quantized, on a Pixel 11 Pro.
- EmbeddingGemma 2 可以用于商业用途吗?可以。它是在 Apache 2.0 许可证下发布的。
- EmbeddingGemma 2 需要多少内存?Google 报告称,在 Pixel 11 Pro 上,仅用于文本时活跃 RAM 约为 191MB,全模态使用时为 567MB(量化后)。
Check out the Model Weights on HF and Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
查看 HF 上的模型权重和技术细节。所有功劳都归于该项目的研究人员。此外,欢迎在 Twitter 上关注我们,别忘了加入我们拥有 150k+ 成员的 ML SubReddit 并订阅我们的新闻通讯。等等!你在 Telegram 上吗?现在你也可以在 Telegram 上加入我们。
[Sponsored] The web is the one API most agents are missing. Databases, calendars and repos have APIs. The open web mostly doesn’t. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free.
[赞助] Web 是大多数智能体缺失的一个 API。数据库、日历和代码仓库都有 API。而开放的网络大多没有。TinyFish MCP 服务器为任何 MCP 客户端提供四个工具:TinySearch、TinyFetch(将完整页面转换为包含 JavaScript 的 Markdown)、用于登录和表单的 TinyBrowser,以及用于多步骤任务的 TinyAgent。搜索和抓取功能免费。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力