Google DeepMind has released EmbeddingGemma 2, an open multimodal embedding model that maps text, code, images, video, and audio into a single embedding space. It follows last year's text-only EmbeddingGemma, which the company says surpassed 20 million downloads. The new model is built on the Gemma 4 architecture and released under the permissive Apache 2.0 license.
Both sources agree on the core specifications: 740 million parameters, a 768-dimensional output that can be truncated to as few as 128 dimensions, and an 8K token context window. Google DeepMind reports best-in-class scores among sub-1B multimodal embedders on benchmarks such as MTEB Code and MAEB, and says it matches or outperforms larger models on text, vision, and audio tasks. MarkTechPost's coverage repeats these figures without adding conflicting details.
The model is designed for on-device deployment. Its modular design requires as little as 270M parameters for text-only workloads, with optional vision and audio encoders. With quantization, text-only weights use about 191MB of active RAM on a Google Pixel 11 Pro, while the full multimodal model needs roughly 567MB. That makes it practical for local semantic search, cross-modal retrieval, and RAG pipelines that keep data on the device.