RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/Llama 4 Maverick
Model reference

Llama 4 Maverick

This reference covers the gated Maverick instruction checkpoint and separates sibling claims from this artifact. It treats provider output limits as third-party listings.

Llama 4 Maverick is the vendor reference for meta-llama/Llama-4-Maverick-17B-128E-Instruct, as of 2026-08-12. Model scale: 17B active, 128 experts, 400B total; Context length: 1M, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identitymeta-llama/Llama-4-Maverick-17B-128E-Instruct
SourceRetrieved 2026-08-12
IdentityHerd release dateApril 5, 2025
SourceRetrieved 2026-08-12
IdentityMaverick identityLlama 4 Maverick, a 17 billion active parameter model with 128 experts
SourceRetrieved 2026-08-12
IdentityScout siblingLlama 4 Scout, a 17 billion active parameter model with 16 experts and 10M context
SourceRetrieved 2026-08-12
IdentityBehemoth announcement statusThe release announcement said: While we're not yet releasing Llama 4 Behemoth as it is still training.
SourceRetrieved 2026-08-12
ArchitectureModel scale17B active, 128 experts, 400B total
SourceRetrieved 2026-08-12
ArchitectureMultimodal fusionnatively multimodal models with early fusion to seamlessly integrate text and vision tokens
SourceRetrieved 2026-08-12
ArchitectureTraining scalemore than 30 trillion tokens
SourceRetrieved 2026-08-12
ContextContext length1M
SourceRetrieved 2026-08-12
ContextOpenRouter provider output limitsProvider limits range from 8K to 32K.
SourceRetrieved 2026-08-12
ModalitiesInputs and outputtext and image input; text output
SourceRetrieved 2026-08-12
ModalitiesLanguages12 supported languages
SourceRetrieved 2026-08-12
LicenseWeights licenseLlama 4 Community License Agreement
SourceRetrieved 2026-08-12
LicenseMonthly-active-user clauseIf, on the Llama 4 version release date, the monthly active users of the products or services made available by or for Licensee...is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta
SourceRetrieved 2026-08-12
LicenseBranding clauseprominently display 'Built with Llama'... you shall also include 'Llama' at the beginning of any such AI model name.
SourceRetrieved 2026-08-12
PricingOpenRouter listingOpenRouter lists $0.20 input and $0.696 output per 1M tokens.
SourceRetrieved 2026-08-12
AvailabilityRepository accessThe Hugging Face repository is gated; access requires Meta's form, and the raw configuration is not publicly fetchable without access.
SourceRetrieved 2026-08-12
AvailabilityVendor deployment guidanceThe FP8 quantized weights fit on a single H100 DGX host while still maintaining quality.
SourceRetrieved 2026-08-12
AvailabilityGated instruction weightsmeta-llama/Llama-4-Maverick-17B-128E-Instruct
SourceRetrieved 2026-08-12

What the vendor says is new

"the first open-weight natively multimodal models with unprecedented context length support and our first built using a mixture-of-experts (MoE) architecture."

SourceRetrieved 2026-08-12

"Llama 4 Scout dramatically increases the supported context length from 128K in Llama 3 to an industry leading 10 million tokens."

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

BenchmarkVendor-claimed valueSource and date
MMLU Pro80.5
SourceRetrieved 2026-08-12
GPQA Diamond69.8
SourceRetrieved 2026-08-12
LiveCodeBench43.4
SourceRetrieved 2026-08-12
LMArena experimental chat versionan experimental chat version scoring ELO of 1417 on LMArena
SourceRetrieved 2026-08-12

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

vLLM

vLLM v0.8.3 supports the Llama 4 herd.

SourceRetrieved 2026-08-12

SGLang

SGLang v0.4.5 includes Llama 4 support from merged pull request 5092.

SourceRetrieved 2026-08-12

Serving concepts

  • Mixture-of-experts serving->
  • Expert parallelism->
  • Context length->
  • KV cache->
  • Tokens per second->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is Llama 4 Maverick?

As of 2026-08-12, release identity: meta-llama/Llama-4-Maverick-17B-128E-Instruct.

Where are the Llama 4 Maverick weights?

As of 2026-08-12, repository access: The Hugging Face repository is gated; access requires Meta's form, and the raw configuration is not publicly fetchable without access.

Can I run Llama 4 Maverick myself?

As of 2026-08-12, vendor deployment guidance: The FP8 quantized weights fit on a single H100 DGX host while still maintaining quality.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy