Loading...
Google's new multimodal embedding model can run on smartphones
It can process images, audio, and video

Google's new multimodal embedding model can run on smartphones

Oct 07, 2026
11:19 am

What's the story

Google has unveiled EmbeddingGemma 2, a new version of its multimodal embedding model. The best part? It's small enough to run on a smartphone. The latest release expands the capabilities of the original model beyond text to include images, audio, and video as well. This means an app built on this new model could match a voice memo with the corresponding moment in a video without sending any data off the phone.

Enhanced capabilities

It has more than double the parameters of its predecessor

The new model is built on the Gemma 4 architecture, which Google released in April. It has more than double the parameters of its predecessor, boasting 740 million parameters.

Most of this increase comes from the vision and audio encoders, which are not needed by apps that only work with text.

Data management

Developers can use a training technique called Matryoshka Representation Learning

An app using EmbeddingGemma 2 also has to store what the model produces. Each embedding is a list of 768 numbers, with every indexed photo, clip, or document adding one more to a local vector database.

However, developers can use a training technique called Matryoshka Representation Learning to cut those lists down as small as 128 numbers.

This can reduce their space by up to six times while still maintaining about 95% of full quality for image, video, and speech.

ADVERTISEMENT

Benchmark results

EmbeddingGemma 2 scored 78.68 on the code section

EmbeddingGemma 2 scored 78.68 on the code section of the Massive Text Embedding Benchmark, nearly 10 points higher than its predecessor.

Google is targeting this result at developers who build retrieval for coding agents.

The company also claims that it outperforms some specialist models more than twice its size on image, video, and audio tasks.

ADVERTISEMENT

Model access

The model weights are now available from Hugging Face

As EmbeddingGemma 2 shares a text tokenizer and an audio encoder with Gemma 4, an on-device retrieval-augmented generation setup running both requires less memory than two unrelated models would.

Google's AI Edge Foresight meeting app for Mac already runs the pair together.

The model weights are now available from Hugging Face and Google's Kaggle under an Apache 2.0 license, allowing commercial use.

ADVERTISEMENT