A 270-million-parameter embedding model that handles text and code, runs through llama.cpp and ships under Apache 2.0 is the piece I actually want for local search. No API in the loop, and the index stays on my own hardware. The modular part is the smart bit: I load the text encoder and skip the image and audio weights until I need them, and the vectors still land in the same space.
What I’d do: point it at my own repos and judge it there. The code-search jump is Google’s own benchmark, and heise is right that real results depend on how you prepare the codebase and what you ask. I’d also start at 256 dimensions and only go up if retrieval gets sloppy.
The story — Google has released EmbeddingGemma 2, which maps text, code, images, audio and video into one shared vector space for search directly on end devices; its predecessor was text-only. Fully loaded it has 740 million parameters. Output defaults to 768 dimensions and can be cut to 128, and the context window is 8192 tokens. Weights are on Hugging Face and Kaggle. (Source)