One balance for optimization, deploys, and the agent. No subscription. Top-ups stay valid for one year.
Everything you need to optimize and ship, on one balance.
$10 minimum top-up, pay for what you run. Unlocks are lifetime purchase totals.
Managed deploy unlocked with up to 4 replicas and 2 GPUs. Add $150 more to unlock managed deploy capacity for 16 replicas and 4 GPUs.
What's included
New accounts start with $1 free to test the inference API.
Add credits to start, $100Dedicated infrastructure, compliance, and custom volume.
Everything self-serve includes, plus
Model APIs
Price per 1M tokens
Input
$0.13Cached input
$0.0189% hit rate
Output
$0.27Input
$0.05Cached input
$0.05Standard rate
Output
$0.15Input
$2.00Cached input
$0.20Output
$6.00Input
$0.10Cached input
$0.01Output
$0.40Input
$0.60Cached input
$0.0396% hit rate
Output
$1.90Input
$0.05Billing basis
USD, billed on input tokens
Input
$0.01Billing basis
USD, billed on input tokens
Input
$0.05Billing basis
USD, billed on input tokens
Input
$0.10Cached input
$0.0165% hit rate
Output
$0.40Input
$0.036Billing basis
USD, billed on input tokens
Limits and modality-specific capabilities are listed on each model page.
Pay as you go is self-serve, funded by credit top-ups. Enterprise adds private infrastructure, reserved capacity, compliance, and custom terms.
Credit pricing and model-serving infrastructure answer different questions. Estimate GPU fit, rent, paid idle capacity, and API crossover for a specific model and traffic shape.
Can't find what you're looking for? Get in touch
What is RunInfra?
Describe what you want to run. RunInfra picks compatible open models, benchmarks GPUs, tunes the runtime, and gives you a deploy-ready stack.
Describe the goal. RunInfra builds and optimizes the stack.
© 2026 RunInfra. All rights reserved.