vLLM
A merged pull request implements Gemma 4 architecture support for MoE, multimodal input, reasoning, and tool use.
This reference covers the dense flagship checkpoint and keeps family-only capabilities separate. It distinguishes model-vendor facts from third-party endpoint pricing.
Gemma 4 31B is the vendor reference for google/gemma-4-31B-it, as of 2026-08-12. Attention architecture: hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global; Context length: 256K tokens for the 12B, 26B, and 31B models, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | google/gemma-4-31B-it | SourceRetrieved 2026-08-12 |
| Identity | Family release date | April 2, 2026 | SourceRetrieved 2026-08-12 |
| Identity | Flagship variant | 31B Dense; 30.7B parameters | SourceRetrieved 2026-08-12 |
| Identity | Family variants | E2B, E4B, 12B, 26B MoE, and 31B Dense; the 26B MoE has 25.2B total, 3.8B active, 8 active experts, 128 total experts, and 1 shared expert. | SourceRetrieved 2026-08-12 |
| Architecture | Attention architecture | hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global | SourceRetrieved 2026-08-12 |
| Architecture | Sliding windows | 512-1024 tokens | SourceRetrieved 2026-08-12 |
| Context | Context length | 256K tokens for the 12B, 26B, and 31B models | SourceRetrieved 2026-08-12 |
| Context | Smaller-family context | 128K tokens for the E2B and E4B models | SourceRetrieved 2026-08-12 |
| Context | Maximum output | Not stated by Google for this artifact. | SourceRetrieved 2026-08-12 |
| Modalities | Inputs | All models natively process video and images, supporting variable resolutions. | SourceRetrieved 2026-08-12 |
| Modalities | Output | text output only | SourceRetrieved 2026-08-12 |
| Modalities | Agentic workflows | Agentic workflows: Native support for function-calling, structured JSON output | SourceRetrieved 2026-08-12 |
| Modalities | Languages | 140+ languages | SourceRetrieved 2026-08-12 |
| Modalities | Audio scope | Audio input is not listed for the 31B model. | SourceRetrieved 2026-08-12 |
| License | Weights license | released under a commercially permissive Apache 2.0 license | SourceRetrieved 2026-08-12 |
| License | Repository access | Apache 2.0 metadata; ungated repository | SourceRetrieved 2026-08-12 |
| License | Use-policy pointer | The vendor model card links a Prohibited use policy page. | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter listing | $0.08 input and $0.35 output per 1M tokens | SourceRetrieved 2026-08-12 |
| Availability | Repository family | Verified repositories include the 31B instruction and base checkpoints, the 26B MoE instruction checkpoint, and the 12B, E4B, and E2B instruction checkpoints, with additional base and QAT variants. | SourceRetrieved 2026-08-12 |
| Availability | Instruction weights | google/gemma-4-31B-it | SourceRetrieved 2026-08-12 |
| Availability | Base weights | google/gemma-4-31B | SourceRetrieved 2026-08-12 |
| Availability | Mixture instruction weights | google/gemma-4-26B-A4B-it | SourceRetrieved 2026-08-12 |
| Availability | Instruction weights | google/gemma-4-12B-it | SourceRetrieved 2026-08-12 |
| Availability | Efficient instruction weights | google/gemma-4-E4B-it | SourceRetrieved 2026-08-12 |
| Availability | Efficient instruction weights | google/gemma-4-E2B-it | SourceRetrieved 2026-08-12 |
"our most intelligent open models to date. Purpose-built for advanced reasoning"
SourceRetrieved 2026-08-12
"breakthrough capabilities made widely accessible under an Apache 2.0 license"
SourceRetrieved 2026-08-12
"intelligence-per-parameter means achieving frontier-level capabilities with less hardware"
SourceRetrieved 2026-08-12
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| Arena AI text leaderboard | 31B: #3 open model in the world on the industry-standard Arena AI text leaderboard | SourceRetrieved 2026-08-12 |
| MMLU Pro | 85.2% | SourceRetrieved 2026-08-12 |
| AIME 2026, no tools | 89.2% | SourceRetrieved 2026-08-12 |
| Codeforces ELO | 2150 | SourceRetrieved 2026-08-12 |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
A merged pull request implements Gemma 4 architecture support for MoE, multimodal input, reasoning, and tool use.
The Gemma 4 cookbook states that all Gemma 4 models require the Triton attention backend for bidirectional image-token attention.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: google/gemma-4-31B-it.
As of 2026-08-12, repository family: Verified repositories include the 31B instruction and base checkpoints, the 26B MoE instruction checkpoint, and the 12B, E4B, and E2B instruction checkpoints, with additional base and QAT variants.
As of 2026-08-12, repository family: Verified repositories include the 31B instruction and base checkpoints, the 26B MoE instruction checkpoint, and the 12B, E4B, and E2B instruction checkpoints, with additional base and QAT variants.
© 2026 RunInfra. All rights reserved.