Google发布多模态嵌入模型EmbeddingGemma 2
EmbeddingGemma 2: An open, lightweight multimodal embedding model
EmbeddingGemma 2: an open, lightweight multimodal embedding model
EmbeddingGemma 2:一款开源、轻量级的多模态嵌入模型
Oct 06, 2026
2026年10月6日
- x.com
- Copy link
- x.com
- 邮件
- 复制链接
EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space.
EmbeddingGemma 2 是用于设备端多模态嵌入的最强大模型,原生地将文本、图像、音频和视频的组合映射到一个统一的嵌入空间中。
Sahil Dua
Research Engineer, Google DeepMind
研究工程师,Google DeepMind
Henrique Schechter Vera
Research Engineer, Google DeepMind
研究工程师,Google DeepMind
Share
分享
- x.com
- Copy link
- x.com
- 邮件
- 复制链接
Your browser does not support the audio element.
您的浏览器不支持音频元素。
Listen to article
收听文章
[[duration]] minutes
[[duration]] 分钟
This content is generated by Google AI. Generative AI is experimental
此内容由 Google AI 生成。生成式 AI 处于实验阶段
Voice Speed
语音速度
Voice
语音
Speed 0.75X 1X 1.5X 2X
速度 0.75X 1X 1.5X 2X
We introduced EmbeddingGemma last year to provide a lightweight option for high-quality text embeddings, to help your apps organize, search, and connect information directly on consumer hardware. The developer community’s response blew past our expectations. With more than 20 million downloads, builders have used it to power smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.
去年我们推出了 EmbeddingGemma,旨在为高质量文本嵌入提供轻量级选项,帮助您的应用在消费级硬件上直接组织、搜索和关联信息。开发者社区的反馈远超我们的预期。凭借超过 2,000 万次下载量,构建者已利用它来驱动更智能的设备端搜索工具以及以隐私优先的检索增强生成(RAG)管道。
Today, we’re launching EmbeddingGemma 2, expanding beyond text to unify code, images, video, and audio in a shared embedding space. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference. It can help find a specific video clip from a voice memo, or search through hours of audio recordings based on a text query, all processed by a single, natively multimodal model.
今天,我们发布 EmbeddingGemma 2,将能力从文本扩展到统一代码、图像、视频和音频的共享嵌入空间。EmbeddingGemma 2 基于 Gemma 4 架构构建,并在商业友好的 Apache 2.0 许可证下发布,拥有 7.4 亿个参数,使其成为设备端推理的理想选择。它可以帮助您从语音备忘录中查找特定的视频片段,或根据文本查询搜索数小时的音频录音,所有这些均由单一的原生多模态模型处理。
Built from the same technology as Gemini Embedding models, EmbeddingGemma 2 is:
EmbeddingGemma 2 采用与 Gemini Embedding 模型相同的技术构建,具备以下特点:
- Best-in-class for its size: Achieves leading scores among sub-1B multimodal embedders for its size across benchmarks like MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark), while matching or outperforming many larger models across text, vision, and audio tasks.
- Modular by design: Requires as little as 270M parameters for text-only workloads with optional vision (170M) and audio (300M) encoders for full multimodal support.
- Storage-efficient: Using Matryoshka Representation Learning (MRL), developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions. This provides up to 6x storage reduction for local vector databases and memory usage.
- Optimized for on-device performance: Runs efficiently within tight resource constraints. With quantization, on a Google Pixel 11 Pro, EmbeddingGemma 2 requires as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model.
- Extended context ready: Features an 8K token context window (4x larger than EmbeddingGemma 1), allowing it to process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations thereof directly on local hardware.
- 同类最佳尺寸表现:在 MTEB(大规模文本嵌入基准测试)Code 和 MAEB(大规模音频嵌入基准测试)等基准测试中,其作为小于 10 亿参数的多模态嵌入器取得了领先分数,同时在文本、视觉和音频任务中匹配或超越了众多更大规模的模型。
- 模块化设计:针对纯文本工作负载仅需 2.7 亿个参数,并可选配视觉(1.7 亿)和音频(3 亿)编码器以实现完整的多模态支持。
- 存储高效:利用嵌套表示学习(Matryoshka Representation Learning, MRL),开发人员可以将输出向量从 768 维动态截断至 512、256 或 128 维。这可为本地向量数据库和内存使用带来高达 6 倍的存储缩减。
- 针对设备端性能优化:在严格的资源限制内高效运行。经过量化后,在 Google Pixel 11 Pro 上,EmbeddingGemma 2 的纯文本权重仅需约 191MB 活跃 RAM,完整多模态模型仅需约 567MB。
- 支持扩展上下文:具备 8K token 的上下文窗口(是 EmbeddingGemma 1 的四倍大),允许其在本地硬件上直接处理长达 5.5 分钟的音频、29 张图像、58 帧视频或它们的交错组合。
Achieving top-tier quality for code, vision, and audio
实现代码、视觉和音频的一流质量
EmbeddingGemma 2 matches the strong multilingual text performance of EmbeddingGemma while delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68), making it well-suited for local codebase indexing, semantic code search, and coding agent retrieval. Across image, video, documents, and audio, it sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size.
EmbeddingGemma 2 在保持与 EmbeddingGemma 相当的多语言文本性能的同时,在代码性能上实现了显著的 9.92 分提升(在 MTEB Code 基准中,从 68.76 提升至 78.68),使其非常适合用于本地代码库索引、语义代码搜索以及编码代理检索。在图像、视频、文档和音频方面,它在参数量低于 10 亿(sub-1B)的模型中树立了质量/参数比的新标准,甚至在某些指标上超越了体积为其两倍以上的一些专用模型。
Find full evaluation metrics and model information in the EmbeddingGemma 2 model card.
在 EmbeddingGemma 2 模型卡片中查找完整的评估指标和模型信息。
Enabling semantic search, routing, and retrieval, fully on-device
实现语义搜索、路由和检索,完全在设备端运行
EmbeddingGemma 2 brings robust capabilities directly to edge hardware. Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and empowers developers to build cross-modal search and retrieval that works entirely offline.
EmbeddingGemma 2 将强大的能力直接带给边缘硬件。在本地生成嵌入向量有助于确保数据隐私,降低流水线延迟,并赋能开发者构建完全离线的跨模态搜索和检索系统。
When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that understand complex multimodal data. Because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.
当与 Gemma 4 等生成式模型配合使用时,EmbeddingGemma 2 能够启用理解复杂多模态数据的设备端 RAG(检索增强生成)流水线。由于 EmbeddingGemma 2 基于 Gemma 4 构建,并共享其文本 tokenizer 和音频编码器,开发者可以在统一的流水线中同时运行这两个模型,从而降低总的内存占用。
Use text or an image to find the top matches in your media library based on semantic similarity. Try it in Google AI Edge Gallery’s Instant Media Search.
使用文本或图像,根据语义相似度在媒体库中查找最匹配的结果。尝试在 Google AI Edge Gallery 的即时媒体搜索(Instant Media Search)中使用它。
Locate specific moments in video using text or audio queries. Try it in Google AI Edge Gallery’s Video Moments Finder.
使用文本或音频查询定位视频中的特定时刻。尝试在 Google AI Edge Gallery 的视频时刻查找器(Video Moments Finder)中使用它。
Pair EmbeddingGemma 2 for local file retrieval with Gemma 4 for contextual reasoning. Try it in the Google AI Edge Foresight app.
将用于本地文件检索的 EmbeddingGemma 2 与用于上下文推理的 Gemma 4 配对使用。尝试在 Google AI Edge Foresight 应用中使用它。
Create real-time decision engines leveraging multimodal context for classification, routing, and predictive capabilities via the MediaPipe Decision Task API.
通过 MediaPipe Decision Task API,利用多模态上下文创建实时决策引擎,以实现分类、路由和预测功能。
To learn how to build on-device search and RAG systems with LiteRT, read the Google AI Edge blog post.
要了解如何使用 LiteRT 构建设备端搜索和 RAG 系统,请阅读 Google AI Edge 博客文章。
Getting started with EmbeddingGemma 2
开始使用 EmbeddingGemma 2
We worked closely with the following partners to ensure EmbeddingGemma 2 works immediately where you build:
我们与以下合作伙伴紧密合作,以确保 EmbeddingGemma 2 能在您构建的地方立即投入使用:
- Download the models: Find the model weights on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability coming soon. Visit LiteRT Community on Hugging Face for models optimized for on-device.
- On-device deployment: Develop cross-platform apps with Google AI Edge MediaPipe for turnkey embedding, retrieval & decision tasks or LiteRT for custom model integration. Build for the browser with transformers.js or WebGPU.
- Use your favorite development tools: Serve the model efficiently using transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio. Store your embedding vectors with Qdrant.
- Fine-tuning: Follow guidance by Unsloth for how to fine-tune EmbeddingGemma 2 for your use cases.
- 下载模型:在 Hugging Face 和 Kaggle 上找到模型权重,Gemini Enterprise Agent Platform Model Garden 即将上线。访问 Hugging Face 上的 LiteRT Community 以获取针对设备端优化的模型。
- 设备端部署:使用 Google AI Edge MediaPipe 开发跨平台应用,以实现开箱即用的嵌入、检索和决策任务,或使用 LiteRT 进行自定义模型集成。使用 transformers.js 或 WebGPU 为浏览器构建应用。
- 使用您喜爱的开发工具:通过 transformers、sentence-transformers、MLX、vLLM、llama.cpp、SGLang、Ollama 和 LMStudio 高效地部署模型。使用 Qdrant 存储您的嵌入向量。
- 微调:遵循 Unsloth 的指导,了解如何为您的用例微调 EmbeddingGemma 2。
Explore our developer guide, documentation, and guides for inference and fine-tuning.
探索我们的开发者指南、文档以及关于推理和微调的指南。
Posted in:
发布于:
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力