vLLM
Support DeepseekV4 was merged; following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass in v0.23.0. This applies to the open V4 lineage because the 0813 weights are not public.
This reference tracks an API build documented separately from an earlier open lineage. It keeps current API facts distinct from lineage details and treats unpublished weights as unavailable.
DeepSeek-V4-Pro-0813 is the vendor reference for DeepSeek-V4-Pro-0813, as of 2026-08-12. Open lineage scale: 1.6T parameters (49B activated); Context length: Context Length: 1M, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | DeepSeek-V4-Pro-0813 | SourceRetrieved 2026-08-12 |
| Identity | API model id | deepseek-v4-pro | SourceRetrieved 2026-08-12 |
| Identity | Availability | GA of DeepSeek V4 Pro; appeared 2026-08-12 | SourceRetrieved 2026-08-12 |
| Architecture | Open lineage scale | 1.6T parameters (49B activated) | SourceRetrieved 2026-08-12 |
| Architecture | Open lineage configuration | 384 routed experts, 6 experts per token, 1 shared expert, 61 layers, 128 attention heads, hidden 7168, vocab 129280, DeepseekV4ForCausalLM | SourceRetrieved 2026-08-12 |
| Architecture | Attention | Novel Attention: Token-wise compression + DSA (DeepSeek Sparse Attention) | SourceRetrieved 2026-08-12 |
| Architecture | Current-build architecture disclosure | The vendor has not stated whether the 0813 build changes the April open-weights architecture. | SourceRetrieved 2026-08-12 |
| Context | Context length | Context Length: 1M | SourceRetrieved 2026-08-12 |
| Context | Maximum output | 384K | SourceRetrieved 2026-08-12 |
| Context | Open lineage position limit | max_position_embeddings 1048576 | SourceRetrieved 2026-08-12 |
| Modalities | Modality | text input and output | SourceRetrieved 2026-08-12 |
| Modalities | Thinking support | thinking and non-thinking modes supported | SourceRetrieved 2026-08-12 |
| Modalities | Open lineage mode names | Non-think, Think High, Think Max | SourceRetrieved 2026-08-12 |
| Modalities | Vision | Not stated by the vendor. | SourceRetrieved 2026-08-12 |
| License | April open lineage | MIT | SourceRetrieved 2026-08-12 |
| Pricing | Input, cache hit | $0.003625 per 1M tokens | SourceRetrieved 2026-08-12 |
| Pricing | Input, cache miss | $0.435 per 1M tokens | SourceRetrieved 2026-08-12 |
| Pricing | Output | $0.87 per 1M tokens | SourceRetrieved 2026-08-12 |
| Availability | Current-build public weights | No public weights for DeepSeek-V4-Pro-0813. | SourceRetrieved 2026-08-12 |
| Availability | Open lineage | DeepSeek-V4-Pro open weights are available. | SourceRetrieved 2026-08-12 |
| Availability | Change log | No 0813 entry was present in the DeepSeek change log. | SourceRetrieved 2026-08-12 |
| Availability | Open lineage weights | deepseek-ai/DeepSeek-V4-Pro | SourceRetrieved 2026-08-12 |
| Availability | Base weights | deepseek-ai/DeepSeek-V4-Pro-Base | SourceRetrieved 2026-08-12 |
| Availability | Speculative decoding draft variant | deepseek-ai/DeepSeek-V4-Pro-DSpark | SourceRetrieved 2026-08-12 |
"The `deepseek-v4-flash` model has been updated to DeepSeek-V4-Flash-0731, and the `deepseek-v4-pro` model has been updated to DeepSeek-V4-Pro-0813."
SourceRetrieved 2026-08-12
"Novel Attention: Token-wise compression + DSA (DeepSeek Sparse Attention)"
SourceRetrieved 2026-08-12
"In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2."
SourceRetrieved 2026-08-12
"We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected."
SourceRetrieved 2026-08-12
These results are vendor-claimed, not independently measured by RunInfra.
No vendor-claimed benchmark results for this exact build were present in the cited material as of 2026-08-12.
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
Support DeepseekV4 was merged; following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass in v0.23.0. This applies to the open V4 lineage because the 0813 weights are not public.
Day-0 support is documented for the open DeepSeek-V4 lineage. This applies to the open lineage because the 0813 weights are not public.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: DeepSeek-V4-Pro-0813.
As of 2026-08-12, input, cache hit: $0.003625 per 1M tokens. As of 2026-08-12, input, cache miss: $0.435 per 1M tokens. As of 2026-08-12, output: $0.87 per 1M tokens.
As of 2026-08-12, current-build public weights: No public weights for DeepSeek-V4-Pro-0813.
© 2026 RunInfra. All rights reserved.