RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/gpt-oss-120b
Model reference

gpt-oss-120b

This reference covers the larger open reasoning checkpoint and its required response format. It keeps self-hosting guidance separate from measured RunInfra evidence.

gpt-oss-120b is the vendor reference for openai/gpt-oss-120b, as of 2026-08-12. Model scale: 116.8B total parameters and 5.1B 'active' parameters per token per forward pass; Context length: 131,072, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identityopenai/gpt-oss-120b
SourceRetrieved 2026-08-12
IdentityPaper submission dateAugust 8, 2025
SourceRetrieved 2026-08-12
IdentitySibling checkpointopenai/gpt-oss-20b
SourceRetrieved 2026-08-12
ArchitectureModel scale116.8B total parameters and 5.1B 'active' parameters per token per forward pass
SourceRetrieved 2026-08-12
ArchitectureMixture configuration36 layers; 128 experts; top-4 routing
SourceRetrieved 2026-08-12
ArchitectureAttention patternbanded window and fully dense patterns alternate, with 128-token windows and grouped-query attention using 64 attention heads and 8 key-value heads
SourceRetrieved 2026-08-12
ArchitectureWeights formatMoE weights quantized to MXFP4 format (4.25 bits per parameter)
SourceRetrieved 2026-08-12
ArchitecturePosition scalingYaRN
SourceRetrieved 2026-08-12
ContextContext length131,072
SourceRetrieved 2026-08-12
ContextMaximum output131,072
SourceRetrieved 2026-08-12
ModalitiesModalityTagged text-generation on the model hub listing.
SourceRetrieved 2026-08-12
ModalitiesReasoning effortConfigurable reasoning effort: Easily adjust the reasoning effort (low, medium, high)
SourceRetrieved 2026-08-12
ModalitiesAgentic capabilitiesAgentic capabilities: Use the models' native capabilities for function calling, web browsing, Python code execution, and Structured Outputs
SourceRetrieved 2026-08-12
LicenseWeights licenseApache 2.0
SourceRetrieved 2026-08-12
LicenseLicense descriptionPermissive Apache 2.0 license: Build freely without copyleft restrictions or patent risk
SourceRetrieved 2026-08-12
PricingOpenRouter listingOpenRouter lists $0.03 per million input tokens and $0.17 per million output tokens.
SourceRetrieved 2026-08-12
AvailabilityRepository accessThe openai/gpt-oss-120b repository is ungated.
SourceRetrieved 2026-08-12
AvailabilityRequired response formatBoth models were trained using our harmony response format and should only be used with this format; otherwise, they will not work correctly.
SourceRetrieved 2026-08-12
AvailabilityVendor deployment guidancefor production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X)
SourceRetrieved 2026-08-12
AvailabilityLarger open reasoning weightsopenai/gpt-oss-120b
SourceRetrieved 2026-08-12
AvailabilitySmaller open reasoning weightsopenai/gpt-oss-20b
SourceRetrieved 2026-08-12

What the vendor says is new

"two open-weight reasoning models that push the frontier of accuracy and inference cost"

SourceRetrieved 2026-08-12

"The models use an efficient mixture-of-expert transformer architecture and are trained using large-scale distillation and reinforcement learning."

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

BenchmarkVendor-claimed valueSource and date
AIME 2025, with tools97.9%
SourceRetrieved 2026-08-12
GPQA Diamond, with tools80.9%
SourceRetrieved 2026-08-12
MMLU90.0%
SourceRetrieved 2026-08-12
SWE-bench Verified62.4%
SourceRetrieved 2026-08-12
Codeforces, with tools2622 Elo
SourceRetrieved 2026-08-12
Paper comparisongpt-oss-120b surpasses OpenAI o3-mini and approaches OpenAI o4-mini accuracy
SourceRetrieved 2026-08-12

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

vLLM

vLLM now supports gpt-oss on NVIDIA Blackwell and Hopper GPUs, as well as AMD MI300x and MI355x GPUs.

SourceRetrieved 2026-08-12

SGLang

The project tracked day-zero gpt-oss support.

SourceRetrieved 2026-08-12

Serving concepts

  • Mixture-of-experts serving->
  • Sliding-window attention->
  • KV cache->
  • Sampling parameters->
  • Tokens per second->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is gpt-oss-120b?

As of 2026-08-12, release identity: openai/gpt-oss-120b.

Where are the gpt-oss-120b weights?

As of 2026-08-12, repository access: The openai/gpt-oss-120b repository is ungated.

Can I run gpt-oss-120b myself?

As of 2026-08-12, vendor deployment guidance: for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X).

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy