· via Hacker News – Front Page (native)
Google's EmbeddingGemma 2 unifies text, image, audio and video embeddings on-device
Google has released EmbeddingGemma 2, a 740-million-parameter open model that maps text, code, images, audio and video into one embedding space and is sized to run on phones and laptops.

Google has released EmbeddingGemma 2, a lightweight multimodal embedding model that maps text, code, images, audio and video into a single shared embedding space and is designed to run entirely on consumer hardware. According to the company's announcement, the model is built on the Gemma 4 architecture, ships under the commercially permissive Apache 2.0 license, and totals 740 million parameters in its full multimodal configuration.
The original EmbeddingGemma arrived last year as a text-only embedder aimed at on-device search and retrieval. Google says it has since passed 20 million downloads, with developers using it for local search tools and privacy-focused retrieval-augmented generation (RAG) pipelines. The second version extends that scope beyond text, pulling code, images, video and audio into the same space.
Benchmarks and architecture
Google says EmbeddingGemma 2 is built from the same technology as its Gemini Embedding models, and claims top results for its size among multimodal embedders under a billion parameters, citing benchmarks including MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark). On code specifically, the company reports a 9.92-point gain on MTEB Code, rising from 68.76 to 78.68 over its predecessor, and suggests the model is a fit for local codebase indexing, semantic code search and coding-agent retrieval.
The model is modular by construction: text-only workloads need roughly 270 million parameters, while optional vision (170 million) and audio (300 million) encoders can be attached for full multimodal coverage. It also applies Matryoshka Representation Learning, which lets developers truncate the default 768-dimensional output vectors down to 512, 256 or 128 dimensions. Google says this can cut storage and memory use for local vector databases by up to 6x.
Sized for phones and laptops
EmbeddingGemma 2 supports an 8K-token context window, four times larger than EmbeddingGemma 1. Google says that capacity covers up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations of them, processed on local hardware.
The memory figures are aimed squarely at handsets: with quantization on a Google Pixel 11 Pro, the model needs as little as roughly 191MB of active RAM for text-only weights and about 567MB for the full multimodal model.
Ecosystem and availability
Google frames the core benefit of local embeddings as privacy and latency, since data never leaves the device and cross-modal search can operate offline. Paired with a generative model such as Gemma 4, the company says EmbeddingGemma 2 supports fully on-device RAG pipelines, and because the two models share a text tokenizer and audio encoder, running them together reduces the combined memory footprint.
Model weights are available on Hugging Face and Kaggle, with the Gemini Enterprise Agent Platform Model Garden listed as coming soon; Google also points to the LiteRT Community on Hugging Face for on-device-optimized variants. Deployment options include MediaPipe and LiteRT for cross-platform apps, and transformers.js with WebGPU for the browser. The model can be served through transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LMStudio, with Qdrant named for vector storage and Unsloth providing fine-tuning guidance.
Why it matters
Embeddings are the connective tissue of modern search and RAG systems, and a single model that embeds text, images, audio and video in one space lets a plain text query surface a video clip or a voice memo without maintaining separate per-modality pipelines. Doing that at 740 million parameters, and a few hundred megabytes of RAM, moves retrieval work from server farms onto the device itself, which matters for privacy-sensitive applications and offline use. The Apache 2.0 license removes commercial friction for adoption, and Google's claim of beating some models more than twice its size points to the broader trend of shrinking, efficient embedders as on-device AI becomes the default expectation. One caveat: the performance figures come from Google's own announcement, so independent benchmarking will show whether the claims hold up outside vendor testing.
- #embeddings
- #on-device-ai
- #multimodal
- #gemma