RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/DeepSeek-V4-Pro-0813
Model reference

DeepSeek-V4-Pro-0813

This reference tracks an API build documented separately from an earlier open lineage. It keeps current API facts distinct from lineage details and treats unpublished weights as unavailable.

DeepSeek-V4-Pro-0813 is the vendor reference for DeepSeek-V4-Pro-0813, as of 2026-08-12. Open lineage scale: 1.6T parameters (49B activated); Context length: Context Length: 1M, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identityDeepSeek-V4-Pro-0813
SourceRetrieved 2026-08-12
IdentityAPI model iddeepseek-v4-pro
SourceRetrieved 2026-08-12
IdentityAvailabilityGA of DeepSeek V4 Pro; appeared 2026-08-12
SourceRetrieved 2026-08-12
ArchitectureOpen lineage scale1.6T parameters (49B activated)
SourceRetrieved 2026-08-12
ArchitectureOpen lineage configuration384 routed experts, 6 experts per token, 1 shared expert, 61 layers, 128 attention heads, hidden 7168, vocab 129280, DeepseekV4ForCausalLM
SourceRetrieved 2026-08-12
ArchitectureAttentionNovel Attention: Token-wise compression + DSA (DeepSeek Sparse Attention)
SourceRetrieved 2026-08-12
ArchitectureCurrent-build architecture disclosureThe vendor has not stated whether the 0813 build changes the April open-weights architecture.
SourceRetrieved 2026-08-12
ContextContext lengthContext Length: 1M
SourceRetrieved 2026-08-12
ContextMaximum output384K
SourceRetrieved 2026-08-12
ContextOpen lineage position limitmax_position_embeddings 1048576
SourceRetrieved 2026-08-12
ModalitiesModalitytext input and output
SourceRetrieved 2026-08-12
ModalitiesThinking supportthinking and non-thinking modes supported
SourceRetrieved 2026-08-12
ModalitiesOpen lineage mode namesNon-think, Think High, Think Max
SourceRetrieved 2026-08-12
ModalitiesVisionNot stated by the vendor.
SourceRetrieved 2026-08-12
LicenseApril open lineageMIT
SourceRetrieved 2026-08-12
PricingInput, cache hit$0.003625 per 1M tokens
SourceRetrieved 2026-08-12
PricingInput, cache miss$0.435 per 1M tokens
SourceRetrieved 2026-08-12
PricingOutput$0.87 per 1M tokens
SourceRetrieved 2026-08-12
AvailabilityCurrent-build public weightsNo public weights for DeepSeek-V4-Pro-0813.
SourceRetrieved 2026-08-12
AvailabilityOpen lineageDeepSeek-V4-Pro open weights are available.
SourceRetrieved 2026-08-12
AvailabilityChange logNo 0813 entry was present in the DeepSeek change log.
SourceRetrieved 2026-08-12
AvailabilityOpen lineage weightsdeepseek-ai/DeepSeek-V4-Pro
SourceRetrieved 2026-08-12
AvailabilityBase weightsdeepseek-ai/DeepSeek-V4-Pro-Base
SourceRetrieved 2026-08-12
AvailabilitySpeculative decoding draft variantdeepseek-ai/DeepSeek-V4-Pro-DSpark
SourceRetrieved 2026-08-12

What the vendor says is new

"The `deepseek-v4-flash` model has been updated to DeepSeek-V4-Flash-0731, and the `deepseek-v4-pro` model has been updated to DeepSeek-V4-Pro-0813."

SourceRetrieved 2026-08-12

"Novel Attention: Token-wise compression + DSA (DeepSeek Sparse Attention)"

SourceRetrieved 2026-08-12

"In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2."

SourceRetrieved 2026-08-12

"We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected."

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

No vendor-claimed benchmark results for this exact build were present in the cited material as of 2026-08-12.

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

vLLM

Support DeepseekV4 was merged; following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass in v0.23.0. This applies to the open V4 lineage because the 0813 weights are not public.

SourceRetrieved 2026-08-12

SGLang

Day-0 support is documented for the open DeepSeek-V4 lineage. This applies to the open lineage because the 0813 weights are not public.

SourceRetrieved 2026-08-12

Serving concepts

  • Mixture-of-experts serving->
  • Expert parallelism->
  • KV cache->
  • Context length->
  • Speculative decoding->
  • Tokens per second->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is DeepSeek-V4-Pro-0813?

As of 2026-08-12, release identity: DeepSeek-V4-Pro-0813.

What does the DeepSeek-V4-Pro-0813 API cost?

As of 2026-08-12, input, cache hit: $0.003625 per 1M tokens. As of 2026-08-12, input, cache miss: $0.435 per 1M tokens. As of 2026-08-12, output: $0.87 per 1M tokens.

Can I run DeepSeek-V4-Pro-0813 myself?

As of 2026-08-12, current-build public weights: No public weights for DeepSeek-V4-Pro-0813.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy