Google's new multimodal embedding model can run on smartphones
What's the story
Google has unveiled EmbeddingGemma 2, a new version of its multimodal embedding model. The best part? It's small enough to run on a smartphone. The latest release expands the capabilities of the original model beyond text to include images, audio, and video as well. This means an app built on this new model could match a voice memo with the corresponding moment in a video without sending any data off the phone.
Enhanced capabilities
It has more than double the parameters of its predecessor
The new model is built on the Gemma 4 architecture, which Google released in April. It has more than double the parameters of its predecessor, boasting 740 million parameters.
Most of this increase comes from the vision and audio encoders, which are not needed by apps that only work with text.
Data management
Developers can use a training technique called Matryoshka Representation Learning
An app using EmbeddingGemma 2 also has to store what the model produces. Each embedding is a list of 768 numbers, with every indexed photo, clip, or document adding one more to a local vector database.
However, developers can use a training technique called Matryoshka Representation Learning to cut those lists down as small as 128 numbers.
This can reduce their space by up to six times while still maintaining about 95% of full quality for image, video, and speech.
Benchmark results
EmbeddingGemma 2 scored 78.68 on the code section
EmbeddingGemma 2 scored 78.68 on the code section of the Massive Text Embedding Benchmark, nearly 10 points higher than its predecessor.
Google is targeting this result at developers who build retrieval for coding agents.
The company also claims that it outperforms some specialist models more than twice its size on image, video, and audio tasks.
Model access
The model weights are now available from Hugging Face
As EmbeddingGemma 2 shares a text tokenizer and an audio encoder with Gemma 4, an on-device retrieval-augmented generation setup running both requires less memory than two unrelated models would.
Google's AI Edge Foresight meeting app for Mac already runs the pair together.
The model weights are now available from Hugging Face and Google's Kaggle under an Apache 2.0 license, allowing commercial use.