Google发布端侧多模态EmbeddingGemma 2
Google dropped EmbeddingGemma 2 for on-device multimodal AI, under an Apache 2.0…
端侧多模态嵌入模型的新突破,模块化设计和低资源占用对移动端开发者极具参考价值,建议关注其在隐私计算场景的应用潜力。
Google dropped EmbeddingGemma 2 for on-device multimodal AI, under an Apache 2.0 license
Google 发布了适用于端侧多模态 AI 的 EmbeddingGemma 2,采用 Apache 2.0 许可证
> puts text, code, images, audio and video into 1 searchable space on phones. gives phones a missing piece: a way to understand and search your own stuff without sending it to a server.
> 将文本、代码、图像、音频和视频统一放入手机上的一个可搜索空间中。 为手机补齐了缺失的一环:一种无需将数据发送至服务器即可理解和搜索个人内容的方式。
> 740M parameters, uses the Gemma 4 architecture.
> 拥有 7.4 亿参数,采用 Gemma 4 架构。
> Its parts are modular, so a text-only app needs just 270M parameters, while a 170M vision encoder and a 300M audio encoder load only when needed.
> 其组件模块化,因此纯文本应用仅需 2.7 亿参数,而视觉编码器和音频编码器(分别为 1.7 亿和 3 亿参数)仅在需要时加载。
> On a Pixel 11 Pro, quantized text weights take about 191MB of active RAM, and the full multimodal model takes about 567MB.
> 在 Pixel 11 Pro 上,量化后的文本权重占用约 191MB 活跃内存,完整的多模态模型占用约 567MB。
> The context window grows 4x to 8K tokens, enough for roughly 5.5 minutes of audio, 29 images or 58 video frames in 1 input.
> 上下文窗口扩大 4 倍至 8K tokens,足以在单次输入中容纳约 5.5 分钟的音频、29 张图像或 58 个视频帧。
> Code search gained most, with the MTEB Code score rising from 68.76 to 78.68 while multilingual text scores held level.
> 代码搜索能力提升最显著,MTEB Code 得分从 68.76 升至 78.68,同时多语言文本得分保持不变。
> Google also claims top sub-1B results on audio and vision benchmarks and wins over some specialist models twice its size,
> Google 还声称其在音频和视觉基准测试中取得了低于 10 亿参数的最佳结果,并击败了一些规模为其两倍的专用模型,
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力