vLLM
A merged pull request enables Qwen3.8 on additional hardware, implying in-tree model support; no release note names it yet.
This reference covers the open text artifact, not its hosted multimodal sibling. It separates artifact facts from hosted service claims and leaves absent pricing unstated.
Qwen3.8-2.4T-A95B is the vendor reference for Qwen/Qwen3.8-2.4T-A95B, as of 2026-08-12. Scale: 2.4T in total and 95B activated; Context: 262,144 natively and extensible up to 1,010,000 tokens, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | Qwen/Qwen3.8-2.4T-A95B | SourceRetrieved 2026-08-12 |
| Identity | FP8 variant | Qwen/Qwen3.8-2.4T-A95B-FP8 | SourceRetrieved 2026-08-12 |
| Identity | Artifact scope | Qwen3.8 Max is the vendor's hosted multimodal API sibling; this page covers only the open-weights text model, and no Max API fact about vision, video, or API pricing applies to this artifact. | SourceRetrieved 2026-08-12 |
| Identity | Repository timing | in the week before 2026-08-12 | SourceRetrieved 2026-08-12 |
| Architecture | Scale | 2.4T in total and 95B activated | SourceRetrieved 2026-08-12 |
| Architecture | Experts | 512 experts; 10 Routed + 1 Shared activated | SourceRetrieved 2026-08-12 |
| Architecture | Layer pattern | 92 layers in the pattern 23 x (3 x (Gated DeltaNet -> MoE) -> 1 x (Gated Attention -> MoE)) | SourceRetrieved 2026-08-12 |
| Architecture | Hybrid linear attention | Gated DeltaNet: 128 linear attention V heads, 16 QK; Gated Attention: 64 Q heads, 4 KV heads, head dim 256; hidden dim 8192 | SourceRetrieved 2026-08-12 |
| Context | Context | 262,144 natively and extensible up to 1,010,000 tokens | SourceRetrieved 2026-08-12 |
| Context | Reasoning content budget | Reasoning Content: 262,144 tokens | SourceRetrieved 2026-08-12 |
| Context | Final response budget | Final Response: 131,072 tokens | SourceRetrieved 2026-08-12 |
| Modalities | Modality | text only for THIS artifact | SourceRetrieved 2026-08-12 |
| Modalities | Thinking mode | requires thinking mode for all interactions | SourceRetrieved 2026-08-12 |
| Modalities | Reasoning format | reasoning in <think> tags | SourceRetrieved 2026-08-12 |
| Modalities | Thinking control | Flexible Thinking Control via reasoning_effort levels xhigh (default), medium, low | SourceRetrieved 2026-08-12 |
| Modalities | Agentic support | tool calling and agentic use supported | SourceRetrieved 2026-08-12 |
| Modalities | Recommended sampling | temperature 1.0, top_p 0.95, top_k 20 | SourceRetrieved 2026-08-12 |
| License | License | Qwen3.8-Max License | SourceRetrieved 2026-08-12 |
| License | License distinction | Not Apache 2.0. | SourceRetrieved 2026-08-12 |
| License | Attribution clause | above >100,000,000 monthly active users or US$ 20,000,000 ... monthly revenue, respective model name must be prominently displayed on the user interface | SourceRetrieved 2026-08-12 |
| License | Model service clause | If the licensee or any of its affiliates conducts a Model as a Service or AI Work Assistant business, and the aggregate revenue ... exceeds US$50,000,000 ... during any consecutive twelve (12) months, the licensee shall obtain a separate license from Qwen before Using the Software or its derivative works for any commercial purpose. | SourceRetrieved 2026-08-12 |
| Pricing | Open artifact pricing | Not published for the open artifact. | SourceRetrieved 2026-08-12 |
| Availability | Reference weights | Qwen/Qwen3.8-2.4T-A95B | SourceRetrieved 2026-08-12 |
| Availability | FP8 weights | Qwen/Qwen3.8-2.4T-A95B-FP8 | SourceRetrieved 2026-08-12 |
| Availability | Vendor-recommended engines | SGLang, vLLM | SourceRetrieved 2026-08-12 |
| Availability | Reference weights | Qwen/Qwen3.8-2.4T-A95B | SourceRetrieved 2026-08-12 |
| Availability | Quantized weights | Qwen/Qwen3.8-2.4T-A95B-FP8 | SourceRetrieved 2026-08-12 |
"the most capable generation in the Qwen open-model family to date"
SourceRetrieved 2026-08-12
"the first time we will open-source the weights of a Qwen-Max-class model"
SourceRetrieved 2026-08-12
"Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains"
SourceRetrieved 2026-08-12
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| PaperBench | 93.0 | SourceRetrieved 2026-08-12 |
| Terminal Bench 2.1 | 86.6 | SourceRetrieved 2026-08-12 |
| GPQA Diamond | 92.6 | SourceRetrieved 2026-08-12 |
| SWE-bench Pro | 67.7 | SourceRetrieved 2026-08-12 |
| HLE | 43.6 | SourceRetrieved 2026-08-12 |
| FrontierSWE | 73.5 | SourceRetrieved 2026-08-12 |
| WideSearch | 81.9 | SourceRetrieved 2026-08-12 |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
A merged pull request enables Qwen3.8 on additional hardware, implying in-tree model support; no release note names it yet.
Merged documentation adds a Qwen3.8 cookbook and serving configs while the core support pull request was still open.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: Qwen/Qwen3.8-2.4T-A95B.
As of 2026-08-12, reference weights: Qwen/Qwen3.8-2.4T-A95B.
As of 2026-08-12, vendor-recommended engines: SGLang, vLLM.
© 2026 RunInfra. All rights reserved.