EmbeddingGemma 2: an open, lightweight multimodal embedding model
EmbeddingGemma 2: an open, lightweight multimodal embedding model
EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space.
Research Engineer, Google DeepMind
Your browser does not support the audio element.
We introduced EmbeddingGemma last year to provide a lightweight option for high-quality text embeddings, to help your apps organize, search, and connect information directly on consumer hardware. The developer community’s response blew past our expectations. With more than 20 million downloads, builders have used it to power smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.
Today, we’re launching EmbeddingGemma 2, expanding beyond text to unify code, images, video, and audio in a shared embedding space. Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference. It can help find a specific video clip from a voice memo, or search through hours of audio recordings based on a text query, all processed by a single, natively multimodal model.
Built from the same technology as Gemini Embedding models, EmbeddingGemma 2 is:
Achieving top-tier quality for code, vision, and audio
EmbeddingGemma 2 matches the strong multilingual text performance of EmbeddingGemma while delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68), making it well-suited for local codebase indexing, semantic code search, and coding agent retrieval. Across image, video, documents, and audio, it sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size.
Find full evaluation metrics and model information in the EmbeddingGemma 2 model card.
Enabling semantic search, routing, and retrieval, fully on-device
EmbeddingGemma 2 brings robust capabilities directly to edge hardware. Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and empowers developers to build cross-modal search and retrieval that works entirely offline.
When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that understand complex multimodal data. Because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.
To learn how to build on-device search and RAG systems with LiteRT, read the Google AI Edge blog post.
Getting started with EmbeddingGemma 2
We worked closely with the following partners to ensure EmbeddingGemma 2 works immediately where you build:
Explore our developer guide, documentation, and guides for inference and fine-tuning.
Related Stories
AI News
Artificial Intelligence in Drug Discovery Market Tops $7.4 billion by 2030, 26% CAGR; Asia
37 minutes ago
AI News
The White House considers anyone who uses the term "artificial intelligence" an enemy
37 minutes ago
AI News
Fox News AI Newsletter: 7 key midterm states put tech boom in focus
1 hour ago
AI News
New meeting-room speakers spread sound across 20
2 hours ago
AI News
Trump draws a crucial line between ‘artificial intelligence’ and ‘super intelligence’
2 hours ago
AI News
Trump: Anyone using the term artificial intelligence is 'the enemy'
2 hours ago
AI News
Trump declares anyone who uses the term Artificial Intelligence to be ‘The Enemy!’
2 hours ago
AI News
Anthropic’s ‘Claude-led’ CRISPR
3 hours ago