Offline is the only kind of search that works at anchor, and the first EmbeddingGemma only did text. EmbeddingGemma 2 puts text, code, images, audio and video into one embedding space, Apache 2.0, at 740 million parameters.
I’d start with the 270M text-only build for local codebase indexing, since MTEB Code jumped from 68.76 to 78.68, serve it through Ollama or llama.cpp, and truncate vectors to 256 dimensions to keep the vector database small. The audio encoder is another 300M; I’d add it when I have recordings worth searching, not before.
The story — Google launched EmbeddingGemma 2, an open multimodal embedding model built on the Gemma 4 architecture under Apache 2.0. It has an 8K token context window and supports truncating vectors from 768 down to 128 dimensions. Quantized on a Pixel 11 Pro, it needs about 191MB of RAM text-only and 567MB fully multimodal. Weights are on Hugging Face and Kaggle. (Source)