RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/DeepSeek-V4-Flash-0731
Model reference

DeepSeek-V4-Flash-0731

This reference restores an archived open checkpoint without implying a current measured package. A future published package with the same identity will replace this view automatically.

DeepSeek-V4-Flash-0731 is the vendor reference for deepseek-ai/DeepSeek-V4-Flash-0731, as of 2026-08-12. Release architecture: DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained.; Context length: 1M, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identityDeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.
SourceRetrieved 2026-08-12
IdentityAPI model iddeepseek-v4-flash
SourceRetrieved 2026-08-12
IdentityRelease date2026-07-31
SourceRetrieved 2026-08-12
ArchitectureRelease architectureDeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained.
SourceRetrieved 2026-08-12
ArchitectureVendor-stated scaleTotal 284B, Activated 13B
SourceRetrieved 2026-08-12
ArchitectureRepository counterThe Hugging Face repository size badge reads 304B from the safetensors count, while the vendor-stated model total is 284B.
SourceRetrieved 2026-08-12
ArchitectureConfiguration43 layers, 64 attention heads, 256 routed experts, 1 shared expert, 6 experts per token, max_position_embeddings 1048576, bfloat16 with fp8 e4m3
SourceRetrieved 2026-08-12
ContextContext length1M
SourceRetrieved 2026-08-12
ContextMaximum outputMAX OUTPUT: MAXIMUM: 384K
SourceRetrieved 2026-08-12
ModalitiesModalitytext
SourceRetrieved 2026-08-12
ModalitiesReasoning effortThe reasoning_effort parameter now supports three levels - low, high, and max - which control how much deliberation the model spends before answering.
SourceRetrieved 2026-08-12
ModalitiesTool callingTool calling is supported.
SourceRetrieved 2026-08-12
LicenseWeights licenseMIT; Copyright (c) 2023 DeepSeek
SourceRetrieved 2026-08-12
PricingDeepSeek API input cache hit$0.0028 per 1M tokens
SourceRetrieved 2026-08-12
PricingDeepSeek API input cache miss$0.14 per 1M tokens
SourceRetrieved 2026-08-12
PricingDeepSeek API output$0.28 per 1M tokens
SourceRetrieved 2026-08-12
PricingDeepSeek pricing warningWe plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected.
SourceRetrieved 2026-08-12
PricingOpenRouter listing$0.08 input and $0.25 output per 1M tokens
SourceRetrieved 2026-08-12
AvailabilityOpen repositorydeepseek-ai/DeepSeek-V4-Flash-0731 exists with about 1.05M downloads at retrieval.
SourceRetrieved 2026-08-12
AvailabilityOpen lineage repositoriesDeepSeek-V4-Flash preview, DeepSeek-V4-Flash-Base, and DeepSeek-V4-Flash-DSpark repositories exist.
SourceRetrieved 2026-08-12
AvailabilityDated-build variantsNo -0731-Base or -0731-DSpark repository is listed for the dated build.
SourceRetrieved 2026-08-12
AvailabilityVendor deployment guidanceThe model card provides vLLM expert-parallel serving and SGLang DSPARK speculative-serving commands.
SourceRetrieved 2026-08-12
AvailabilityDated instruction weightsdeepseek-ai/DeepSeek-V4-Flash-0731
SourceRetrieved 2026-08-12
AvailabilityPreview instruction weightsdeepseek-ai/DeepSeek-V4-Flash
SourceRetrieved 2026-08-12
AvailabilityBase weightsdeepseek-ai/DeepSeek-V4-Flash-Base
SourceRetrieved 2026-08-12
AvailabilitySpeculative decoding draft variantdeepseek-ai/DeepSeek-V4-Flash-DSpark
SourceRetrieved 2026-08-12

What the vendor says is new

"DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained."

SourceRetrieved 2026-08-12

"The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex."

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

BenchmarkVendor-claimed valueSource and date
Terminal Bench 2.182.7
SourceRetrieved 2026-08-12
NL2Repo54.2
SourceRetrieved 2026-08-12
Cybergym76.7
SourceRetrieved 2026-08-12
DeepSWE54.4
SourceRetrieved 2026-08-12
Toolathlon verified70.3
SourceRetrieved 2026-08-12

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

vLLM

The DeepSeek V4 Rebased pull request adding initial DeepSeek V4 support was merged on April 27, 2026.

SourceRetrieved 2026-08-12

SGLang

Day-zero support covers the open DeepSeek V4 family.

SourceRetrieved 2026-08-12

Serving concepts

  • Mixture-of-experts serving->
  • Speculative decoding->
  • KV cache->
  • Context length->
  • Tokens per second->
  • Expert parallelism->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is DeepSeek-V4-Flash-0731?

As of 2026-08-12, release identity: DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.

Where are the DeepSeek-V4-Flash-0731 weights?

As of 2026-08-12, open repository: deepseek-ai/DeepSeek-V4-Flash-0731 exists with about 1.05M downloads at retrieval.

Can I run DeepSeek-V4-Flash-0731 myself?

As of 2026-08-12, vendor deployment guidance: The model card provides vLLM expert-parallel serving and SGLang DSPARK speculative-serving commands.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy