RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Model Library

Compare hosted model availability, prices, measured speed, and supported capabilities.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
  1. Model Library
  2. /
  3. DeepSeek
  4. /
  5. DeepSeek V4 Pro

DeepSeek V4 Pro

deepseek-ai/DeepSeek-V4-Pro-0813
Get API keyView docs

DeepSeek V4 Pro is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4-Pro-0813 at $0.60 per 1M input tokens and $1.90 per 1M output tokens. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions.

Pricing

USD, pay per token

per 1M input tokens
$0.60
per 1M output tokens
$1.90

Measured performance

Output speed

207output tokens per second, model only

Time to first token

520milliseconds to first token

Access

Confirm how your client reaches this model.

Provider
DeepSeek
API compatibility
OpenAI-compatible chat completions
Accepted input
Text only
Availability
Available
Data retention
Your prompts are never stored and never used for training. Billing keeps token counts, cost, timing, and a request id. An Idempotency-Key you send holds that response body for 24 hours.

Capacity

Check the limits your workload must fit.

Context window
1,048,576 tokens
Maximum request size
3.5 MB per request
Maximum generated output
32,768 tokens
Streaming time to first token limit
60 seconds
Maximum request duration
240 seconds

Capabilities

See which request modes the API supports.

Tool calling
Supported
JSON mode
Supported
Streaming
Supported
View Full Spec
Gateway compatibility
OpenAI-compatible chat completions for compatible clients and gateways
OpenRouter Provider Monitor
Provider metadata published as ready for discovery. Not proof of a live OpenRouter listing or callability.
Prefix caching
Automatic prefix caching runs on every replica. Requests are spread across replicas without session affinity, so a repeated prompt is not guaranteed to reach the cache holding it. Cached input is billed at the standard input rate. Responses do not report cached token counts.
Upstream model
View model

Code examples

Set RUNINFRA_GATEWAY_KEY to your workspace API key before using an example.

Trust and provenance

Verify the company and operating credentials behind this API.

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
  • Y CombinatorBacked by Y Combinator.
  • NVIDIA InceptionMember of NVIDIA Inception.
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy