· via dev.to (home feed)
EmbeddingGemma 2 is a 740M-parameter multimodal embedder built for on-device search
Google DeepMind has released EmbeddingGemma 2, a 740M-parameter open model that maps text, code, images, video and audio into one shared vector space, with modular encoders aimed at on-device search and RAG.

Google DeepMind has released EmbeddingGemma 2, a 740-million-parameter open model that projects text, code, images, video and audio into a single shared 768-dimensional vector space. According to a technical write-up on dev.to, the model is built on the Gemma 4 architecture, ships under an Apache 2.0 license, and is deliberately sized to run on consumer hardware rather than requiring datacenter-class infrastructure.
One vector space for five modalities
The unification is the central pitch. Because all five input types are embedded into the same coordinate space, a developer can run genuinely cross-modal retrieval from one index: find a moment in a video library using a plain text query, or pull up a code snippet described in a voice memo. According to the dev.to article, which draws on coverage from Unite.AI and MarkTechPost, this removes the need to stitch together separate embedders per modality when building multimodal search or retrieval-augmented generation pipelines.
The model is also positioned as an enabler for privacy-preserving applications, since retrieval that runs entirely on a phone or laptop means the underlying user data never has to be sent to a remote service.
Modular encoders keep the memory footprint small
The most consequential design decision for builders is that the 740M parameters are not one monolithic block. The write-up breaks the model into independent components: a 270M-parameter text encoder, a 170M-parameter vision encoder and a 300M-parameter audio encoder, all loadable from a single checkpoint.
That means an application loads only what it needs. A text-and-code RAG pipeline on a low-memory device can pull in just the 270M text component, while a photo-organisation tool would load the combined 440M text-plus-vision configuration. The reported numbers illustrate the payoff: on a Pixel 11 Pro, the text-only variant occupies roughly 191MB of active memory, while the full multimodal setup needs about 567MB. For anyone shipping to mid-range phones, that gap is the difference between feasible and not.
Benchmark gains and dimension trade-offs
On quality, the article reports that EmbeddingGemma 2 improves code retrieval over its predecessor by 9.92 points on the MTEB Code benchmark, a meaningful jump for a model in this size class.
The model also supports Matryoshka Representation Learning, which lets developers truncate the 768-dimensional vectors to 512, 256 or 128 dimensions to cut storage and compute costs. The trade-offs are modality-specific, according to the write-up: for text-only workloads, dropping to 256 dimensions costs little on multilingual benchmarks, but image tasks degrade more visibly at 128 dimensions, so the smallest size is recommended mainly for text.
Why it matters
Embeddings are the connective tissue of modern search and RAG systems, and until now multimodal embedding has largely meant either large server-side models or gluing together per-modality embedders, which makes cross-modal queries awkward. An Apache-2.0-licensed model small enough for a phone, with one shared space across text, code, images, video and audio, changes what an individual developer can ship: semantic search over a user's photos, screenshots, recordings and documents that runs locally and keeps data on the device.
The modular encoder loading and Matryoshka truncation are the details that make it practical, giving engineers direct control over the memory, storage and quality balance on constrained hardware. One caveat worth noting: the benchmark figures and device measurements cited here come from initial coverage of the release rather than independent evaluation, so teams evaluating the model should validate performance on their own data.
- #embeddings
- #on-device-ai
- #google-deepmind
- #multimodal
- #rag