Skip to content
Text, image, audio, video and code cards flow into an open embedding model, then cluster as points in a shared vector space for a search query.

BlogIndustry News

Multimodal Embeddings Go Local: What EmbeddingGemma 2 Means

Google's open EmbeddingGemma 2 puts text, image, audio and video search on your own hardware. Here's what it changes and how to pilot it.

Google released EmbeddingGemma 2 on October 6: an open model that turns text, code, images, video and audio into vectors in one shared space. That matters, because multimodal embeddings have mostly been a cloud API you pay for per call and send your data to. This one has 740 million parameters, ships under Apache 2.0, and is small enough to run on a phone.

For a business owner, the real question isn't which benchmark it tops. It's whether you can finally search product photos, support call recordings, PDFs and code with one index, without shipping that data to a third party. For a lot of teams, the answer just moved to yes, with a few caveats.

TL;DR#

  • EmbeddingGemma 2 is a 740M-parameter open model that maps text, code, images, video and audio into one 768-dimension vector space, licensed Apache 2.0.
  • It runs locally. Google reports about 191MB of RAM for text-only use and about 567MB for the full multimodal model on a Pixel 11 Pro with quantization.
  • The opportunity is one search index across your messy files. The catch: no safety tuning, an 8K-token budget per call, and benchmark numbers that come from Google itself.

What Google released#

EmbeddingGemma 2 is the follow-up to the first EmbeddingGemma, which Google says passed 20 million downloads. It's built on the Gemma 4 architecture and uses the same technology as Google's Gemini Embedding models. The model card on Hugging Face gives the details.

SpecEmbeddingGemma 2
Parameters740M total (270M text core, 170M vision, 300M audio)
InputsText, code, images, video, audio, mixed in one call
Output size768 dimensions, can be cut to 512, 256 or 128
Context8,192 tokens shared across all inputs
LicenseApache 2.0 (commercial use allowed)
Runs ontransformers, sentence-transformers, vLLM, llama.cpp, Ollama, MLX, LiteRT, transformers.js

The 8K window covers roughly 29 images, 58 video frames or about 5.5 minutes of audio per call. On code retrieval, Google reports the MTEB Code score rose from 68.76 to 78.68 compared with the first version. Weights are on Hugging Face and Kaggle now, and Google says availability in its enterprise Model Garden is coming soon.

What multimodal embeddings actually do#

An embedding model turns a piece of content into a long list of numbers. Content with similar meaning ends up with similar numbers, so "find things like this" becomes a fast math problem instead of a keyword match. That's the engine behind semantic search and most retrieval-augmented generation (RAG) setups.

A multimodal embedding model does this across formats in the same space. A photo of a frame with a cracked corner, the support email saying "the corner arrived damaged" and a 30-second voicemail about it should all land close together. You can search with any of them and find the others.

Why this release changes the math#

Multimodal search isn't new. Cloud options like Amazon Nova Multimodal Embeddings already exist. What's different here is the combination: open weights, a commercial license, all five input types in one model, and a footprint small enough for a laptop, an edge box or a modest server. As The Next Web points out, removing the network step also removes the data-transfer question that worries compliance teams.

Technical breakdown and trade-offs#

Load only what you need#

The encoders are modular. Text-only loads the 270M core, text plus images is 440M, and the full model is 740M. You don't pay memory for audio support you never use.

Smaller vectors, smaller bills#

Matryoshka Representation Learning lets you truncate vectors from 768 to 256 or 128 dimensions. Google says that cuts storage up to 6x. The model card is honest about the cost: quality stays close to full at 256, but at 128 the overall multimodal score (MMEB) falls to 45.65 and the multilingual text score drops from 61.36 to 57.89. So 256 is a sensible default, and 128 is something you test, not assume.

The fine print engineers will hit#

  • Precision: run it in bfloat16 or float32. The model card warns it returns NaN or degraded embeddings in float16.
  • Task prefixes: text queries work best with short prefixes like task: search result | query:. Skip them and quality drops.
  • No safety tuning: it's a pre-trained model with no output moderation. Guardrails are your job.

Cloud API or self-hosted?#

FactorCloud embedding APISelf-hosted open model
Setup timeHoursDays
Data leaves your systemsYesNo
Cost shapePer call, grows with volumeFixed hardware, flat
UpgradesVendor decidesYou decide, and you re-index
Ops burdenLowYou own uptime and monitoring

High-volume or sensitive data leans self-hosted. A small pilot often starts faster on an API.

What it means for businesses#

The useful way to think about this is: which of your files are invisible today because they aren't text?

  • E-commerce: customers search by uploading a photo, and your catalog returns visually similar products, even when the product titles are inconsistent.
  • Support: an agent pastes a screenshot or a call snippet and gets the five most similar past tickets, with how they were resolved.
  • Compliance and ops: scanned delivery notes, inspection photos and recorded calls become searchable without sending them to an outside service.
  • Engineering: coding agents retrieve the right internal code faster. Google calls out local codebase indexing as a target use case.

The privacy angle is the one that moves budgets. If legal has blocked AI search because data would leave your environment, a local open-weight model removes that objection. You still need access controls, retention rules and logging, but the conversation gets much shorter.

How to act on it: a practical checklist#

  1. Pick one painful search. "Find past orders with damage photos like this one" beats "AI search for everything."
  2. Gather 200 to 500 real examples with known correct matches. This is your test set.
  3. Run a two-day spike. Embed the set at 768 and 256 dimensions and measure how often the right answer lands in the top five.
  4. Compare against your current option, whether that's keyword search or a cloud embedding API.
  5. Decide where it runs: on-device, on your own server, or in your cloud account.
  6. Add guardrails: permission filters on results, logging, and a human check for anything customer-facing.
  7. Plan for re-indexing. Changing embedding models means re-embedding everything, so record which model and version made each vector.

A minimal pipeline looks like this:

files (pdf, jpg, mp3, mp4) -> EmbeddingGemma 2 -> vectors (256d)
                                                      |
user query (text / photo / audio) -> embed -> vector DB top-k -> filters -> results

And the first few lines in Python:

from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2")
q = model.encode("damaged corner on frame", prompt_name="SearchQuery")

How MagicMakers Lab approaches this#

We treat a multimodal embeddings project as an integration job, not a model demo. The model is the easy part. The work is connecting it to where your files already live (your store, helpdesk, ERP or shared drives), keeping permissions intact and measuring whether results are actually better. On GeoVerdant we took satellite-data analysis from weeks to minutes by building around the data the team already had, and the same approach applies here through our AI Integration & Automation work.

Key takeaways#

  • EmbeddingGemma 2 brings text, image, video, audio and code embeddings into one open, commercially licensed model.
  • It's small enough to run locally, which takes data transfer off the table for sensitive files.
  • 256 dimensions is a safe starting point. Test 128 before relying on it.
  • No safety tuning means filters, permissions and logging are on you.
  • Start with one high-value search problem and a real test set, not a platform rebuild.

FAQ#

What are multimodal embeddings?#

Multimodal embeddings are numeric representations of content where text, images, audio and video share the same vector space. Items with similar meaning sit close together regardless of format, so you can search photos with a sentence or find documents with an audio clip. They power cross-format semantic search, recommendations and retrieval for AI assistants.

What is an embedding model used for in business?#

An embedding model powers search that understands meaning rather than exact words. Common uses include product search, finding similar support tickets, grouping customer feedback, detecting duplicates and feeding the right documents to an AI assistant (RAG). It doesn't write answers itself. It finds the most relevant material so people or other models can act on it.

Can EmbeddingGemma 2 be used commercially?#

Yes. Google released it under the Apache 2.0 license, which allows commercial use, modification and self-hosting. You still need to follow the Gemma Prohibited Use Policy referenced in the model card. Because the model has no safety tuning or output moderation, you're responsible for application-level safeguards such as permission filtering and logging.

Is a local embedding model better than a cloud embedding API?#

It depends on volume and data sensitivity. Local models keep data inside your systems and turn per-call costs into fixed hardware costs, which suits sensitive or high-volume workloads. Cloud APIs are faster to start with and need less operations work. Many teams pilot on an API, then move high-volume or regulated workloads to a self-hosted model.

If you've got product photos, call recordings or scanned documents that nobody can search today, that's usually the best place to start. We'll look at your stack and tell you plainly whether a multimodal search pilot is worth it. Book a free audit.

Sources#