RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/Qwen3.8-2.4T-A95B
Model reference

Qwen3.8-2.4T-A95B

This reference covers the open text artifact, not its hosted multimodal sibling. It separates artifact facts from hosted service claims and leaves absent pricing unstated.

Qwen3.8-2.4T-A95B is the vendor reference for Qwen/Qwen3.8-2.4T-A95B, as of 2026-08-12. Scale: 2.4T in total and 95B activated; Context: 262,144 natively and extensible up to 1,010,000 tokens, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identityQwen/Qwen3.8-2.4T-A95B
SourceRetrieved 2026-08-12
IdentityFP8 variantQwen/Qwen3.8-2.4T-A95B-FP8
SourceRetrieved 2026-08-12
IdentityArtifact scopeQwen3.8 Max is the vendor's hosted multimodal API sibling; this page covers only the open-weights text model, and no Max API fact about vision, video, or API pricing applies to this artifact.
SourceRetrieved 2026-08-12
IdentityRepository timingin the week before 2026-08-12
SourceRetrieved 2026-08-12
ArchitectureScale2.4T in total and 95B activated
SourceRetrieved 2026-08-12
ArchitectureExperts512 experts; 10 Routed + 1 Shared activated
SourceRetrieved 2026-08-12
ArchitectureLayer pattern92 layers in the pattern 23 x (3 x (Gated DeltaNet -> MoE) -> 1 x (Gated Attention -> MoE))
SourceRetrieved 2026-08-12
ArchitectureHybrid linear attentionGated DeltaNet: 128 linear attention V heads, 16 QK; Gated Attention: 64 Q heads, 4 KV heads, head dim 256; hidden dim 8192
SourceRetrieved 2026-08-12
ContextContext262,144 natively and extensible up to 1,010,000 tokens
SourceRetrieved 2026-08-12
ContextReasoning content budgetReasoning Content: 262,144 tokens
SourceRetrieved 2026-08-12
ContextFinal response budgetFinal Response: 131,072 tokens
SourceRetrieved 2026-08-12
ModalitiesModalitytext only for THIS artifact
SourceRetrieved 2026-08-12
ModalitiesThinking moderequires thinking mode for all interactions
SourceRetrieved 2026-08-12
ModalitiesReasoning formatreasoning in <think> tags
SourceRetrieved 2026-08-12
ModalitiesThinking controlFlexible Thinking Control via reasoning_effort levels xhigh (default), medium, low
SourceRetrieved 2026-08-12
ModalitiesAgentic supporttool calling and agentic use supported
SourceRetrieved 2026-08-12
ModalitiesRecommended samplingtemperature 1.0, top_p 0.95, top_k 20
SourceRetrieved 2026-08-12
LicenseLicenseQwen3.8-Max License
SourceRetrieved 2026-08-12
LicenseLicense distinctionNot Apache 2.0.
SourceRetrieved 2026-08-12
LicenseAttribution clauseabove >100,000,000 monthly active users or US$ 20,000,000 ... monthly revenue, respective model name must be prominently displayed on the user interface
SourceRetrieved 2026-08-12
LicenseModel service clauseIf the licensee or any of its affiliates conducts a Model as a Service or AI Work Assistant business, and the aggregate revenue ... exceeds US$50,000,000 ... during any consecutive twelve (12) months, the licensee shall obtain a separate license from Qwen before Using the Software or its derivative works for any commercial purpose.
SourceRetrieved 2026-08-12
PricingOpen artifact pricingNot published for the open artifact.
SourceRetrieved 2026-08-12
AvailabilityReference weightsQwen/Qwen3.8-2.4T-A95B
SourceRetrieved 2026-08-12
AvailabilityFP8 weightsQwen/Qwen3.8-2.4T-A95B-FP8
SourceRetrieved 2026-08-12
AvailabilityVendor-recommended enginesSGLang, vLLM
SourceRetrieved 2026-08-12
AvailabilityReference weightsQwen/Qwen3.8-2.4T-A95B
SourceRetrieved 2026-08-12
AvailabilityQuantized weightsQwen/Qwen3.8-2.4T-A95B-FP8
SourceRetrieved 2026-08-12

What the vendor says is new

"the most capable generation in the Qwen open-model family to date"

SourceRetrieved 2026-08-12

"the first time we will open-source the weights of a Qwen-Max-class model"

SourceRetrieved 2026-08-12

"Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains"

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

BenchmarkVendor-claimed valueSource and date
PaperBench93.0
SourceRetrieved 2026-08-12
Terminal Bench 2.186.6
SourceRetrieved 2026-08-12
GPQA Diamond92.6
SourceRetrieved 2026-08-12
SWE-bench Pro67.7
SourceRetrieved 2026-08-12
HLE43.6
SourceRetrieved 2026-08-12
FrontierSWE73.5
SourceRetrieved 2026-08-12
WideSearch81.9
SourceRetrieved 2026-08-12

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

vLLM

A merged pull request enables Qwen3.8 on additional hardware, implying in-tree model support; no release note names it yet.

SourceRetrieved 2026-08-12

SGLang

Merged documentation adds a Qwen3.8 cookbook and serving configs while the core support pull request was still open.

SourceRetrieved 2026-08-12

Serving concepts

  • Mixture-of-experts serving->
  • Expert parallelism->
  • Context length->
  • Sampling parameters->
  • KV cache->
  • Tokens per second->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is Qwen3.8-2.4T-A95B?

As of 2026-08-12, release identity: Qwen/Qwen3.8-2.4T-A95B.

Where are the Qwen3.8-2.4T-A95B weights?

As of 2026-08-12, reference weights: Qwen/Qwen3.8-2.4T-A95B.

Can I run Qwen3.8-2.4T-A95B myself?

As of 2026-08-12, vendor-recommended engines: SGLang, vLLM.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy