Google releases EmbeddingGemma 2 for on-device multimodal search
The downloadable 740-million-parameter model puts text, code, images, video and audio in one embedding space.
One embedding space for several media types
Google DeepMind released EmbeddingGemma 2 on October 6, with downloadable weights on Hugging Face under Apache 2.0. The 740-million-parameter model turns text and code, images, video and audio into vectors in a shared 768-dimensional space. That lets a search system compare a text query with other media without using a separate embedding model for each input type. Google says the model is designed for local use on phones and laptops, including offline search and retrieval applications.
The model card describes a 270-million-parameter text component and optional vision and audio encoders. It also specifies an 8,192-token shared context window and allows shorter output vectors for smaller indexes. Those savings come with tradeoffs: the card says 128-dimensional vectors notably reduce multimodal quality, and warns against float16 inference because it may produce invalid or silently degraded results. Google’s performance claims come from its published evaluations, not independent tests of every device or workload. Developers can inspect the weights and model card, but should measure retrieval quality and memory requirements in their own applications.