# Google's EmbeddingGemma 2 lifts MTEB Code by 9.92 points, runs on-device

From The Forward Pass daily issue, October 7, 2026 (https://theforwardpass.net/archive/daily/2026-10-07). Source: https://theforwardpass.net/archive/daily/2026-10-07/googles-embeddinggemma-2-lifts-mteb-code-by-9-92-points-runs-on-device

> This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

**Top News** · 275 HN points

EmbeddingGemma 2 maps text, images, audio, video and code into one shared space. It is built for local search, retrieval and RAG.

Here's what you can build with it:
- Its MTEB Code score rises from 68.76 to 78.68 over EmbeddingGemma.
- The full model has 740M parameters, and text-only use needs as little as 270M.
- On a Pixel 11 Pro, quantized RAM is about 191MB for text-only or 567MB multimodal.
- Matryoshka vectors shrink from 768 to 128 dimensions for up to sixfold storage savings.
- An 8K-token context window, four times EmbeddingGemma 1, and Apache 2.0 licensing.

Pair it with Gemma 4 for local RAG, since they share a text tokenizer and audio encoder.

One catch: Google gives no numerical speed results.

Try it: download weights from Hugging Face or Kaggle, deploy with MediaPipe or LiteRT, or run in the browser with transformers.js. It also serves via sentence-transformers, vLLM, llama.cpp and Ollama.

Sources: [blog.google](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/) · [deepmind.google](https://deepmind.google/blog/embeddinggemma-2-an-open-lightweight-multimodal-embedding-model/)
