vLLM
vLLM now supports gpt-oss on NVIDIA Blackwell and Hopper GPUs, as well as AMD MI300x and MI355x GPUs.
This reference covers the larger open reasoning checkpoint and its required response format. It keeps self-hosting guidance separate from measured RunInfra evidence.
gpt-oss-120b is the vendor reference for openai/gpt-oss-120b, as of 2026-08-12. Model scale: 116.8B total parameters and 5.1B 'active' parameters per token per forward pass; Context length: 131,072, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | openai/gpt-oss-120b | SourceRetrieved 2026-08-12 |
| Identity | Paper submission date | August 8, 2025 | SourceRetrieved 2026-08-12 |
| Identity | Sibling checkpoint | openai/gpt-oss-20b | SourceRetrieved 2026-08-12 |
| Architecture | Model scale | 116.8B total parameters and 5.1B 'active' parameters per token per forward pass | SourceRetrieved 2026-08-12 |
| Architecture | Mixture configuration | 36 layers; 128 experts; top-4 routing | SourceRetrieved 2026-08-12 |
| Architecture | Attention pattern | banded window and fully dense patterns alternate, with 128-token windows and grouped-query attention using 64 attention heads and 8 key-value heads | SourceRetrieved 2026-08-12 |
| Architecture | Weights format | MoE weights quantized to MXFP4 format (4.25 bits per parameter) | SourceRetrieved 2026-08-12 |
| Architecture | Position scaling | YaRN | SourceRetrieved 2026-08-12 |
| Context | Context length | 131,072 | SourceRetrieved 2026-08-12 |
| Context | Maximum output | 131,072 | SourceRetrieved 2026-08-12 |
| Modalities | Modality | Tagged text-generation on the model hub listing. | SourceRetrieved 2026-08-12 |
| Modalities | Reasoning effort | Configurable reasoning effort: Easily adjust the reasoning effort (low, medium, high) | SourceRetrieved 2026-08-12 |
| Modalities | Agentic capabilities | Agentic capabilities: Use the models' native capabilities for function calling, web browsing, Python code execution, and Structured Outputs | SourceRetrieved 2026-08-12 |
| License | Weights license | Apache 2.0 | SourceRetrieved 2026-08-12 |
| License | License description | Permissive Apache 2.0 license: Build freely without copyleft restrictions or patent risk | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter listing | OpenRouter lists $0.03 per million input tokens and $0.17 per million output tokens. | SourceRetrieved 2026-08-12 |
| Availability | Repository access | The openai/gpt-oss-120b repository is ungated. | SourceRetrieved 2026-08-12 |
| Availability | Required response format | Both models were trained using our harmony response format and should only be used with this format; otherwise, they will not work correctly. | SourceRetrieved 2026-08-12 |
| Availability | Vendor deployment guidance | for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) | SourceRetrieved 2026-08-12 |
| Availability | Larger open reasoning weights | openai/gpt-oss-120b | SourceRetrieved 2026-08-12 |
| Availability | Smaller open reasoning weights | openai/gpt-oss-20b | SourceRetrieved 2026-08-12 |
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| AIME 2025, with tools | 97.9% | SourceRetrieved 2026-08-12 |
| GPQA Diamond, with tools | 80.9% | SourceRetrieved 2026-08-12 |
| MMLU | 90.0% | SourceRetrieved 2026-08-12 |
| SWE-bench Verified | 62.4% | SourceRetrieved 2026-08-12 |
| Codeforces, with tools | 2622 Elo | SourceRetrieved 2026-08-12 |
| Paper comparison | gpt-oss-120b surpasses OpenAI o3-mini and approaches OpenAI o4-mini accuracy | SourceRetrieved 2026-08-12 |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: openai/gpt-oss-120b.
As of 2026-08-12, repository access: The openai/gpt-oss-120b repository is ungated.
As of 2026-08-12, vendor deployment guidance: for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X).
© 2026 RunInfra. All rights reserved.