Models / EmbeddingGemma 2

The local-RAG default just moved. Google’s EmbeddingGemma 2 is a 744M-parameter embedding model, Apache 2.0, that turns text (and, per the family line, image and audio inputs) into a vector a search index can score. 21,148 direct downloads and 1,087 likes in days; community builds (unsloth GGUF at 29,700 downloads, an ONNX export, a LiteRT-LLM port for phone GPUs) landed within the week.

Why embeddings matter more than chat models do locally. A retrieval stack runs the embedding model on EVERY query and EVERY document - thousands of calls where a chat model runs one. At 744M params the fp16 build is about 1.5GB and a 4-bit one is about 0.4GB, so the whole index-brain fits in RAM that would otherwise hold one layer of a chat model, and the 48MB-per-1,000-token KV cost barely registers. That is the difference between an index that re-embeds weekly and one that re-embeds nightly.

Where it sits. Google’s original embeddinggemma (300M) proved phone-scale retrieval at 3.6M downloads; version 2 brings the 8k-context class, and the open-weights Apache license puts it next to BGE-M3 and GTE as the default pick for a self-hosted RAG stack that wants a current-generation embedding model without a server.

embeddings rag retrieval local
Parameters
740M
Context
262k
License
apache 2.0
Developer
Google
Origin
🇺🇸 USA
Released
Oct 2026

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

Q4_K_M
0.4GB 1.5GB min 2.0GB rec
Balanced - the usual local sweet spot
FP16
1.5GB 2.0GB min 3.0GB rec
Full quality, largest

The reference hardware

Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth
Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig Q4_K_M FP16
4x H100 80GB (320GB) fast 12449.3t/s fast 4355.8t/s
NVIDIA DGX Station 748GB fast 7432.4t/s fast 2600.5t/s
8x RTX 3090 rack (192GB) fast 6958.2t/s fast 2434.6t/s
4x RTX 5090 (128GB) fast 6659.5t/s fast 2330.0t/s
AMD Instinct MI300X (192GB) fast 4947.0t/s fast 1730.9t/s
4x RTX 4090 (96GB) fast 3746.0t/s fast 1310.6t/s
2x RTX 5090 (64GB) fast 3329.7t/s fast 1165.0t/s
2x RTX 3090 (48GB) fast 1739.6t/s fast 608.6t/s
Single RTX 5090 (32GB) fast 1664.9t/s fast 582.5t/s
RTX PRO 6000 Blackwell (96GB) fast 1664.9t/s fast 582.5t/s
Mac Studio M4 Ultra 192GB fast 1106.8t/s fast 387.2t/s
Mac Studio M4 Ultra 512GB fast 1106.8t/s fast 387.2t/s
Single RTX 4090 (24GB) fast 936.5t/s fast 327.7t/s
MacBook Pro M5 Max 128GB fast 622.3t/s fast 217.7t/s
Single GTX 1080 Ti (11GB) fast 449.7t/s fast 157.3t/s
Dual EPYC 9004 + 768GB DDR5-4800 fast 428.1t/s fast 149.8t/s
DGX Spark 128GB unified fast 253.6t/s fast 88.7t/s
Ryzen AI Max+ 395 128GB fast 237.8t/s fast 83.2t/s
Jetson AGX Orin 64GB fast 190.3t/s fast 66.6t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 fast 190.3t/s fast 66.6t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 fast 142.7t/s fast 49.9t/s
NVIDIA Jetson Orin NX 16GB fast 95.1t/s fast 33.3t/s

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options

Q4_K_M community
0.4GB dl 1.5GB min 2.0GB rec
REC RAM vs largest quant
0.4GB q4 weights; phone-and-laptop class
FP16 official
1.5GB dl 2.0GB min 3.0GB rec
REC RAM vs largest quant
1.5GB fp16 weights + KV 48MB/1k; an embedding model's cache is per-call and freed

Or run it in the cloud

No per-token API provider pricing tracked for EmbeddingGemma 2 yet. For flagship list prices, see the calculator.

See who runs Google in production →

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...