deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Google DeepMind open-sources EmbeddingGemma 2 for fully local multimodal RAG

Google DeepMind's EmbeddingGemma 2 maps text, code, images, video and audio into one 768-dimensional vector space under Apache 2.0, enabling retrieval pipelines that never leave the device.

Google DeepMind open-sources EmbeddingGemma 2 for fully local multimodal RAG

Google DeepMind has released EmbeddingGemma 2, an open-weight multimodal embedding model that places text, code, images, video and audio into a single 768-dimensional vector space. According to a dev.to post, the model arrived on October 6, 2026 under the Apache 2.0 license, and its main appeal is architectural: retrieval-augmented generation (RAG) pipelines in which every embedding and every similarity search happens on the device holding the data, with no call out to a third-party embedding API.

Modular encoders, one shared space

EmbeddingGemma 2 is not a single monolithic network but a set of modular encoders that produce vectors in a compatible space. The dev.to write-up lists four configurations: a 270M-parameter build for text and code, aimed at documentation and knowledge bases; a 440M text-plus-vision variant for images, diagrams and video frames; a 570M text-plus-audio variant for recordings and speech; and a 740M full multimodal configuration covering everything.

Because all modalities share one representation, a plain-text query can be scored directly against an image or an audio clip, without generating captions or transcripts as an intermediate step. One point the post stresses: this is an embedding model, not a generative LLM. It finds relevant content but does not write answers, so a fully on-device RAG stack still needs a generative model that also runs locally.

Working with it in Python

Google documents the model for sentence-transformers 6.1.0 or newer. The dev.to walkthrough shows loading google/embeddinggemma-2 with the vision and audio encoders disabled for a lighter text-only setup, encoding documents and queries with named prompts ("Document" and "SearchQuery"), and ranking passages by similarity. Media is handled by passing a dictionary keyed by modality, such as an image file or a meeting recording, whose vectors can then be compared with a text query. The first run downloads the weights; after that, encoding can work offline from cached resources.

Truncatable vectors and storage trade-offs

EmbeddingGemma 2 supports Matryoshka Representation Learning, so embeddings can be shortened from 768 to 512, 256 or 128 dimensions. The post offers rough arithmetic: one million float32 vectors occupy about 3.07 GB at full dimensionality versus roughly 1.02 GB at 256, excluding metadata and index overhead. Queries and indexed content must use matching dimensions, and truncated vectors should be checked for retrieval quality on real data. Google has published memory figures for optimised configurations on specific devices, but the post cautions these are not universal guarantees for arbitrary Python applications or browsers.

Local is not automatically private

The post is careful not to oversell the privacy angle. Local embeddings reduce certain data transfers but do not replace a complete security design. Deployments still need answers to questions such as whether model downloads, logs, analytics or synchronisation transmit sensitive data, how embeddings and metadata are protected at rest, whether retrieval enforces document-level permissions before showing results, whether retrieved content is treated as untrusted input to blunt prompt injection, and whether the generative step runs on-device or ships selected passages to an API.

On the practical side, long recordings should be indexed as timestamped segments rather than one vector per file, since a whole-video vector cannot reliably locate a particular moment, and scanned PDFs may need OCR alongside image-based retrieval. The post also proposes an evaluation method, noting that no independent benchmarks exist yet: compare remote embeddings, local text-only retrieval and local multimodal retrieval on the same corpus and hardware, using 30-50 questions and metrics such as recall@k, retrieval latency, peak memory and storage, repeated at both 768 and 256 dimensions.

Why it matters

EmbeddingGemma 2 removes two obstacles that have kept multimodal RAG tethered to cloud APIs for many teams: licensing and data flow. An Apache 2.0 release with encoders small enough, between 270M and 740M parameters, to plausibly run on a laptop makes it realistic to search internal documentation, codebases, screenshots and meeting recordings without that content ever leaving the machine. The shared vector space also simplifies architectures that previously stitched together captioning models, transcribers and separate text encoders.

The model remains one component among many. Parsing, chunking, permissions, incremental index updates, evaluation and a separate generator are still the developer's job. But as a building block for privacy-sensitive and offline-capable search, it shifts the central engineering question from which API should receive this data toward which pipeline stages the local device can reasonably run.

  • #embeddings
  • #rag
  • #open-source
  • #multimodal
  • #google-deepmind

Related posts