GOOGLE / 2026
Gemma 4 26B A4B IT
A compact multimodal MoE with 25.2B total and 3.8B active parameters, designed for accessible inference.
DECISION BRIEF
Single-accelerator baseline
The pinned BF16 checkpoint carries 51.6 GB of tensors, clearing the static baseline on one 80 GB accelerator while still requiring runtime validation.
Publicly available checkpoints; verify terms before commercial deployment.
Unknown remains explicit until the exact checkpoint license is verified and sourced.
Architecture values are linked to the publishing organization.
ARCHITECTURE
Compute is sparse. Residency is not.
The pinned BF16 checkpoint carries 51.6 GB of tensors, clearing the static baseline on one 80 GB accelerator while still requiring runtime validation.
Expert topology: 128 / top-8. The active-parameter count approximates token-level compute; it does not determine checkpoint memory, KV-cache demand, expert placement, or interconnect pressure.
Google introduces Gemma 4
google/gemma-4-26B-A4B-it@4d7ae4984b7d · 51.6 GB
NEXT ACTION
Test the workload, not the headline.
Start with checkpoint residency, then layer in context, concurrency, runtime overhead, topology, and resilience.
Build an assurance plan