vLLM
The DeepSeek V4 Rebased pull request adding initial DeepSeek V4 support was merged on April 27, 2026.
This reference restores an archived open checkpoint without implying a current measured package. A future published package with the same identity will replace this view automatically.
DeepSeek-V4-Flash-0731 is the vendor reference for deepseek-ai/DeepSeek-V4-Flash-0731, as of 2026-08-12. Release architecture: DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained.; Context length: 1M, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. | SourceRetrieved 2026-08-12 |
| Identity | API model id | deepseek-v4-flash | SourceRetrieved 2026-08-12 |
| Identity | Release date | 2026-07-31 | SourceRetrieved 2026-08-12 |
| Architecture | Release architecture | DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained. | SourceRetrieved 2026-08-12 |
| Architecture | Vendor-stated scale | Total 284B, Activated 13B | SourceRetrieved 2026-08-12 |
| Architecture | Repository counter | The Hugging Face repository size badge reads 304B from the safetensors count, while the vendor-stated model total is 284B. | SourceRetrieved 2026-08-12 |
| Architecture | Configuration | 43 layers, 64 attention heads, 256 routed experts, 1 shared expert, 6 experts per token, max_position_embeddings 1048576, bfloat16 with fp8 e4m3 | SourceRetrieved 2026-08-12 |
| Context | Context length | 1M | SourceRetrieved 2026-08-12 |
| Context | Maximum output | MAX OUTPUT: MAXIMUM: 384K | SourceRetrieved 2026-08-12 |
| Modalities | Modality | text | SourceRetrieved 2026-08-12 |
| Modalities | Reasoning effort | The reasoning_effort parameter now supports three levels - low, high, and max - which control how much deliberation the model spends before answering. | SourceRetrieved 2026-08-12 |
| Modalities | Tool calling | Tool calling is supported. | SourceRetrieved 2026-08-12 |
| License | Weights license | MIT; Copyright (c) 2023 DeepSeek | SourceRetrieved 2026-08-12 |
| Pricing | DeepSeek API input cache hit | $0.0028 per 1M tokens | SourceRetrieved 2026-08-12 |
| Pricing | DeepSeek API input cache miss | $0.14 per 1M tokens | SourceRetrieved 2026-08-12 |
| Pricing | DeepSeek API output | $0.28 per 1M tokens | SourceRetrieved 2026-08-12 |
| Pricing | DeepSeek pricing warning | We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter listing | $0.08 input and $0.25 output per 1M tokens | SourceRetrieved 2026-08-12 |
| Availability | Open repository | deepseek-ai/DeepSeek-V4-Flash-0731 exists with about 1.05M downloads at retrieval. | SourceRetrieved 2026-08-12 |
| Availability | Open lineage repositories | DeepSeek-V4-Flash preview, DeepSeek-V4-Flash-Base, and DeepSeek-V4-Flash-DSpark repositories exist. | SourceRetrieved 2026-08-12 |
| Availability | Dated-build variants | No -0731-Base or -0731-DSpark repository is listed for the dated build. | SourceRetrieved 2026-08-12 |
| Availability | Vendor deployment guidance | The model card provides vLLM expert-parallel serving and SGLang DSPARK speculative-serving commands. | SourceRetrieved 2026-08-12 |
| Availability | Dated instruction weights | deepseek-ai/DeepSeek-V4-Flash-0731 | SourceRetrieved 2026-08-12 |
| Availability | Preview instruction weights | deepseek-ai/DeepSeek-V4-Flash | SourceRetrieved 2026-08-12 |
| Availability | Base weights | deepseek-ai/DeepSeek-V4-Flash-Base | SourceRetrieved 2026-08-12 |
| Availability | Speculative decoding draft variant | deepseek-ai/DeepSeek-V4-Flash-DSpark | SourceRetrieved 2026-08-12 |
These results are vendor-claimed, not independently measured by RunInfra.
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.
As of 2026-08-12, open repository: deepseek-ai/DeepSeek-V4-Flash-0731 exists with about 1.05M downloads at retrieval.
As of 2026-08-12, vendor deployment guidance: The model card provides vLLM expert-parallel serving and SGLang DSPARK speculative-serving commands.
© 2026 RunInfra. All rights reserved.