RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Start with $10 in credits, pay for what you run.

One balance for optimization, deploys, and the agent. No subscription. Top-ups stay valid for one year.

Pay as you go

Everything you need to optimize and ship, on one balance.

No subscription

$10 minimum top-up, pay for what you run. Unlocks are lifetime purchase totals.

Managed deploy unlocked with up to 4 replicas and 2 GPUs. Add $150 more to unlock managed deploy capacity for 16 replicas and 4 GPUs.

What's included

Quantization with AWQ, GPTQ, and FP8
All standard GPUs (T4, L4, L40S, A100, H100)
Managed deploy with scale-to-zero endpoints, from $50 lifetime
OpenAI-compatible API endpoints
Deployment kit included with every optimization, your exit is free
Agent chat, plans, and benchmarking
One balance, no seats, no per-user fees
Unlimited pipelines and versioning

New accounts start with $1 free to test the inference API.

Add credits to start, $100

Enterprise

Dedicated infrastructure, compliance, and custom volume.

Custom pricing

Everything self-serve includes, plus

Self-hosted and custom-GPU deployment
Audit logs and role-based access control (RBAC)
B200 / H200 GPU access
Custom volume and contract terms
Custom SLAs up to 99.99%
SOC 2 Type II attestation
Dedicated CSM and private Slack
Request a demo

Model APIs

Pay for model tokens, not idle capacity

Price per 1M tokens

Model

Input

Cached input

Output

Model page
DeepSeek V4 Flash

Input

$0.13

Cached input

$0.01

89% hit rate

Output

$0.27
View model→
Nemotron 3.5 Lightning 30B

Input

$0.05

Cached input

$0.05

Standard rate

Output

$0.15
View model→
Qwen3.8 2.4T A95B

Input

$2.00

Cached input

$0.20

Output

$6.00
View model→
Qwen3.8 27B

Input

$0.10

Cached input

$0.01

Output

$0.40
View model→
DeepSeek V4 Pro

Input

$0.60

Cached input

$0.03

96% hit rate

Output

$1.90
View model→
Qwen3 Embedding 8B

Input

$0.05

Billing basis

USD, billed on input tokens

View model→
Qwen3 Embedding 0.6B

Input

$0.01

Billing basis

USD, billed on input tokens

View model→
Qwen3 Reranker 8B

Input

$0.05

Billing basis

USD, billed on input tokens

View model→
Ornith 1.5 35B

Input

$0.10

Cached input

$0.01

65% hit rate

Output

$0.40
View model→
Parakeet TDT 0.6B v3

Input

$0.036

Billing basis

USD, billed on input tokens

View model→

Limits and modality-specific capabilities are listed on each model page.

Compare Pay as you go and Enterprise

Pay as you go is self-serve, funded by credit top-ups. Enterprise adds private infrastructure, reserved capacity, compliance, and custom terms.

Pay as you go
from a $10 top-up
Pricing modelCredit top-ups from $10, pay for what you run
Capability milestones$50: up to 4 replicas and 2 GPUs. $250: up to 16 replicas and 4 GPUs. $10,000: up to 32 replicas and 8 GPUs. Lifetime spend.
BalanceOne balance for optimization, deploys, and inference
SeatsNo per-seat fees
Optimization techniques
Standard GPUs (T4 to H100)
B200 / H200 GPUs
Managed deployUnlocks at $50 lifetime spend
Self-hosted / custom GPUIncluded (deployment kit + BYOC)
OpenAI-compatible API
Audit logs and RBAC
SOC 2 Type II
SLANo contractual SLA
SupportPriority email
Enterprise
Custom
Pricing modelCustom volume and terms
Capability milestonesIncluded under contract
BalanceOne shared balance with custom controls
SeatsNo per-seat fees
Optimization techniques
Standard GPUs (T4 to H100)
B200 / H200 GPUs
Managed deploy
Self-hosted / custom GPUIncluded, plus custom-GPU deployment
OpenAI-compatible API
Audit logs and RBAC
SOC 2 Type II
SLAUp to 99.99%
SupportDedicated CSM and private Slack
Compare Pay as you go and Enterprise

Pay as you go is self-serve, funded by credit top-ups. Enterprise adds private infrastructure, reserved capacity, compliance, and custom terms.

Pay as you go
from a $10 top-up
Enterprise
Custom
Pricing modelCredit top-ups from $10, pay for what you runCustom volume and terms
Capability milestones$50: up to 4 replicas and 2 GPUs. $250: up to 16 replicas and 4 GPUs. $10,000: up to 32 replicas and 8 GPUs. Lifetime spend.Included under contract
BalanceOne balance for optimization, deploys, and inferenceOne shared balance with custom controls
SeatsNo per-seat feesNo per-seat fees
Optimization techniques
Standard GPUs (T4 to H100)
B200 / H200 GPUs
Managed deployUnlocks at $50 lifetime spend
Self-hosted / custom GPUIncluded (deployment kit + BYOC)Included, plus custom-GPU deployment
OpenAI-compatible API
Audit logs and RBAC
SOC 2 Type II
SLANo contractual SLAUp to 99.99%
SupportPriority emailDedicated CSM and private Slack

Credit pricing and model-serving infrastructure answer different questions. Estimate GPU fit, rent, paid idle capacity, and API crossover for a specific model and traffic shape.

Common questions

Can't find what you're looking for? Get in touch

What is RunInfra?

Describe what you want to run. RunInfra picks compatible open models, benchmarks GPUs, tunes the runtime, and gives you a deploy-ready stack.

How do I build my first pipeline?

Type what you want, like 'a support copilot with Whisper and Qwen, tuned for latency.' RunInfra builds and optimizes the pipeline. Chat to refine it, then deploy.

Which AI models are supported?

Vetted Hugging Face models. LLM serving is fully supported end to end. Speech, embedding, vision, and image generation models are supported for optimization and benchmarking, with managed serving in staged rollout. Gated or unsupported models are flagged before you start, not after.

How does GPU kernel optimization work?

RunInfra profiles your model across GPUs, tries quantization, KV cache, serving, and kernel tweaks, and benchmarks the best tradeoff of speed, memory, and cost.

Can I deploy pipelines as APIs?

Yes. Supported pipelines deploy as REST endpoints in one click, and managed deploy unlocks at $50 of lifetime credit purchases. If something isn't deployable yet, RunInfra tells you why instead of shipping a broken endpoint.

How is this different from using closed-source APIs?

Closed APIs hide the model and the infrastructure. With RunInfra you see both, and you benchmark open models against your own latency, throughput, and cost targets. Export the stack or run it in your own cloud to own it outright. Managed hosting is the convenient option, with a free deployment kit as your exit.

Is my data secure?

Encrypted in transit and at rest, on isolated infrastructure. Your inference data never trains anything. Run an exported kit in your own cloud and your inference data stays there. Use our managed endpoints and it passes through the subprocessors listed in our privacy policy. RunInfra is SOC 2 Type II attested.

How does the balance work?

You add credits with one-time top-ups, starting at $10, and spend from one balance, in dollars. Optional auto-recharge adds credits automatically when your balance runs low. Top-ups stay valid for one year, and the same balance covers agent plans, optimization, benchmarking, deploys, and hosted inference.

Deploy your first optimized model, measured before you ship

Describe the goal. RunInfra builds and optimizes the stack.

Start BuildingView Pricing
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

Join the communitySystem status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy