The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

04SEPT2025replayed
one year on
model launchGoogle DeepMind

Google releases EmbeddingGemma, a 308M-parameter open-weights text embedding model for on-device AI

Small enough to run on mobile phones and laptops with under 200MB of RAM, the multilingual model is designed for private RAG and semantic search.

Google today released EmbeddingGemma, a 308-million-parameter open-weights text embedding model built on the Gemma 3 architecture. The model is designed to run directly on edge devices such as mobile phones, laptops, and desktops, using less than 200MB of RAM with quantization. It supports over 100 languages and delivers customizable output dimensions from 768 down to 128 via Matryoshka representation, with a 2K token context window.

Google claims EmbeddingGemma is the highest-ranking open multilingual text embedding model under 500 million parameters on the MTEB benchmark, with inference under 15 milliseconds for 256 tokens on EdgeTPU. The model is engineered for privacy-preserving on-device applications including retrieval-augmented generation and semantic search, operating entirely offline. Weights are available on Hugging Face, Kaggle, and Vertex AI, with integrations already supported across tools such as sentence-transformers, llama.cpp, MLX, Ollama, LiteRT, transformers.js, and LangChain.

“EmbeddingGemma is comparable to popular models nearly twice its size,” the team wrote. The release comes as Google pushes mobile-first AI pipelines, pairing EmbeddingGemma with its on-device Gemma 3n generative model so that both retrieval and generation can run locally, without a network round-trip.

One year later — open only if you can handle spoilers

EmbeddingGemma settled in as a default for on-device and local RAG, still turning up in 2026's roundups of the best small open embedding models and wired into local-AI tooling like Ollama, llama.cpp and sentence-transformers. It never topped the overall embedding leaderboards — larger and proprietary models kept those — but that was never the pitch: the point was useful retrieval that runs on a phone, offline.

Filed

Compiled and edited by .

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy