AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Google announced EmbeddingGemma 2 on Oct. 6, 2026, a 740-million-parameter model designed to map text, code, images, audio and video into a shared embedding space. Google says it supports local multimodal search and retrieval, with an Apache 2.0 license; its benchmark and device-performance claims are company-reported.

Google announced EmbeddingGemma 2 on Oct. 6, describing it as a 740-million-parameter multimodal embedding model that maps text, code, images, audio and video into a shared representation space. Released under the Apache 2.0 license, it is intended to support search and retrieval directly on consumer devices, including queries that connect information across different media.

Google says the model is built on the Gemma 4 architecture and can handle combinations of modalities within an 8,000-token context window. The company says that window can cover up to 5.5 minutes of audio, 29 images or 58 video frames, or interleaved combinations. EmbeddingGemma 2 is designed for uses such as finding a video clip from a voice memo or searching audio recordings with a text query.

The model is modular: Google lists a 270-million-parameter text-only configuration, with optional vision and audio encoders for broader multimodal use. It also uses Matryoshka Representation Learning, allowing developers to reduce output vectors from 768 dimensions to 512, 256 or 128. Google says this can cut storage and memory needs by up to six times, depending on the configuration.

Google reports that, with quantization, the model uses about 191 MB of active RAM for text-only weights and about 567 MB for the full multimodal version on a Pixel 11 Pro. The company also reports a 9.92-point gain over the original EmbeddingGemma on MTEB Code, rising from 68.76 to 78.68. These figures and claims about performance against other models are from Google; independent evaluation details were not provided in the announcement itself.

At a glance
announcementWhen: Announced Oct. 6, 2026
The developmentGoogle released EmbeddingGemma 2, an open-weight multimodal embedding model designed to run on consumer hardware.

Local Search Across Media Types

EmbeddingGemma 2 targets a practical limitation of many search systems: text, images, recordings and video often need separate models or processing steps before they can be searched together. A shared embedding space could let developers retrieve related material across formats, such as locating a video using words spoken in an audio note. If the model performs as Google reports, it may make that kind of search more feasible on phones and other resource-constrained devices.

Running embedding generation locally can also reduce the need to send private content to a remote service and may lower latency or keep some features available offline. Those are potential benefits, not guarantees: privacy depends on an app’s full design, and device speed will vary with hardware and workload. The permissive license and compatibility with several development frameworks could widen access to the model, though adoption and real-world results remain to be seen.

Amazon

multimodal search device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Text Embeddings to Multimodal

Google introduced the first EmbeddingGemma model last year as a smaller option for text embeddings, used to organize information and support semantic search and retrieval-augmented generation. In its announcement, Google said that model had passed 20 million downloads. That figure is a company-reported download count, not a measure of active use or independent adoption.

The new release extends the family to image, audio and video alongside text and code. Google says EmbeddingGemma 2 shares its text tokenizer and audio encoder with Gemma 4, which the company says can help developers run both in one on-device pipeline with a lower combined memory footprint. It is available through Hugging Face and Kaggle; Google says availability in Gemini Enterprise Agent Platform Model Garden is planned for a later date.

““natively mapping combinations of text, images, audio, and video into a unified embedding space.””

— Google, in its Oct. 6, 2026 announcement

Amazon

audio and video search tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Testing Still Needed

The announcement does not include independent benchmark results or a detailed comparison methodology for the claims that EmbeddingGemma 2 leads models below one billion parameters or exceeds some larger specialist models. Google directs readers to the model card for full evaluation metrics, but results can depend on benchmark versions, hardware and evaluation settings. The figures should therefore be read as Google-reported results until independently replicated.

It is also not yet clear how the model’s speed, power use and retrieval quality will vary across devices beyond the Pixel 11 Pro example, or how much memory different combinations of modalities require in typical applications. Google has not specified a date for Model Garden availability. The announcement describes intended uses and technical capacity; it does not establish that every application will gain the same privacy, latency or offline benefits.

Amazon

on-device AI embedding models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Model Access and Developer Tests

Developers can download the weights from Hugging Face and Kaggle and test deployment through Google AI Edge MediaPipe or LiteRT. Google also lists support across tools including transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio, with browser development options using transformers.js or WebGPU. Fine-tuning guidance is available through Unsloth, according to the announcement.

The next useful evidence will come from developers testing the model on different devices and comparing its quality, memory use and speed with alternatives on clearly specified tasks. Google says Gemini Enterprise Agent Platform Model Garden availability is coming soon, but has not announced a date. Independent evaluations and practical deployments should clarify how well the model’s stated multimodal capabilities carry over to everyday local search systems.

Amazon

multimedia search software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is EmbeddingGemma 2?

It is a 740-million-parameter embedding model from Google that represents text, code, images, audio and video in a shared space to support search and retrieval across media.

Can EmbeddingGemma 2 run on a phone?

Google designed it for on-device use and reports about 191 MB of active RAM for quantized text-only weights on a Pixel 11 Pro, and about 567 MB for the full multimodal model on that device. Requirements and performance may differ on other hardware.

What license does the model use?

Google says EmbeddingGemma 2 is released under the Apache 2.0 license, which permits commercial use subject to the license’s terms.

Where can developers get it?

Google says the weights are available on Hugging Face and Kaggle. Availability in Gemini Enterprise Agent Platform Model Garden is planned, with no date specified in the announcement.

Are its benchmark claims independently verified?

The announcement presents benchmark figures and comparisons as Google-reported results. It does not provide independent verification; Google points developers to the model card for evaluation details.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Opus 5.5 Agents Discover Two Room-temperature Magnetic Semiconductor Candidates

Vals AI reports two computationally predicted candidates for room-temperature magnetic semiconductors, but experimental confirmation is still needed.

What Are The Four Capacity Challenges For US Data Centers?

Rymvard’s illustrative scenarios show how grid delays, curtailment, cooling and utility charges can separate reserved power from usable capacity.

The 12 Most Effective AI Productivity Software For 2026

Discover the 12 most effective AI productivity tools for 2026, their features, benefits, and what makes them stand out for different workflows.