vLLM
vLLM v0.8.3 supports the Llama 4 herd.
This reference covers the gated Maverick instruction checkpoint and separates sibling claims from this artifact. It treats provider output limits as third-party listings.
Llama 4 Maverick is the vendor reference for meta-llama/Llama-4-Maverick-17B-128E-Instruct, as of 2026-08-12. Model scale: 17B active, 128 experts, 400B total; Context length: 1M, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | meta-llama/Llama-4-Maverick-17B-128E-Instruct | SourceRetrieved 2026-08-12 |
| Identity | Herd release date | April 5, 2025 | SourceRetrieved 2026-08-12 |
| Identity | Maverick identity | Llama 4 Maverick, a 17 billion active parameter model with 128 experts | SourceRetrieved 2026-08-12 |
| Identity | Scout sibling | Llama 4 Scout, a 17 billion active parameter model with 16 experts and 10M context | SourceRetrieved 2026-08-12 |
| Identity | Behemoth announcement status | The release announcement said: While we're not yet releasing Llama 4 Behemoth as it is still training. | SourceRetrieved 2026-08-12 |
| Architecture | Model scale | 17B active, 128 experts, 400B total | SourceRetrieved 2026-08-12 |
| Architecture | Multimodal fusion | natively multimodal models with early fusion to seamlessly integrate text and vision tokens | SourceRetrieved 2026-08-12 |
| Architecture | Training scale | more than 30 trillion tokens | SourceRetrieved 2026-08-12 |
| Context | Context length | 1M | SourceRetrieved 2026-08-12 |
| Context | OpenRouter provider output limits | Provider limits range from 8K to 32K. | SourceRetrieved 2026-08-12 |
| Modalities | Inputs and output | text and image input; text output | SourceRetrieved 2026-08-12 |
| Modalities | Languages | 12 supported languages | SourceRetrieved 2026-08-12 |
| License | Weights license | Llama 4 Community License Agreement | SourceRetrieved 2026-08-12 |
| License | Monthly-active-user clause | If, on the Llama 4 version release date, the monthly active users of the products or services made available by or for Licensee...is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta | SourceRetrieved 2026-08-12 |
| License | Branding clause | prominently display 'Built with Llama'... you shall also include 'Llama' at the beginning of any such AI model name. | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter listing | OpenRouter lists $0.20 input and $0.696 output per 1M tokens. | SourceRetrieved 2026-08-12 |
| Availability | Repository access | The Hugging Face repository is gated; access requires Meta's form, and the raw configuration is not publicly fetchable without access. | SourceRetrieved 2026-08-12 |
| Availability | Vendor deployment guidance | The FP8 quantized weights fit on a single H100 DGX host while still maintaining quality. | SourceRetrieved 2026-08-12 |
| Availability | Gated instruction weights | meta-llama/Llama-4-Maverick-17B-128E-Instruct | SourceRetrieved 2026-08-12 |
"the first open-weight natively multimodal models with unprecedented context length support and our first built using a mixture-of-experts (MoE) architecture."
SourceRetrieved 2026-08-12
"Llama 4 Scout dramatically increases the supported context length from 128K in Llama 3 to an industry leading 10 million tokens."
SourceRetrieved 2026-08-12
These results are vendor-claimed, not independently measured by RunInfra.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: meta-llama/Llama-4-Maverick-17B-128E-Instruct.
As of 2026-08-12, repository access: The Hugging Face repository is gated; access requires Meta's form, and the raw configuration is not publicly fetchable without access.
As of 2026-08-12, vendor deployment guidance: The FP8 quantized weights fit on a single H100 DGX host while still maintaining quality.
© 2026 RunInfra. All rights reserved.