One balance for optimization, deploys, and the agent. No subscription. Top-ups stay valid for one year.
Hosted Model APIs, per token, on one prepaid balance.
$10 minimum top-up, pay for what you run.
What's included
New accounts start with $1 free to test the Model APIs.
Add credits to start, $100Dedicated infrastructure, compliance, and custom volume.
Everything self-serve includes, plus
Model APIs
Price per 1M tokens
Input
$0.05Cached input
$0.0196.7% cache hit, last 24 hours
Output
$0.15Input
$0.10Cached input
$0.0193.3% cache hit, last 24 hours
Output
$0.40Input
$0.10Cached input
$0.0193.0% cache hit, last 24 hours
Output
$0.40Input
$0.10Cached input
$0.0198.6% cache hit, last 24 hours
Output
$0.40Limits and modality-specific capabilities are listed on each model page.
Pay as you go is self-serve, funded by credit top-ups. Enterprise adds private infrastructure, reserved capacity, compliance, and custom terms.
Credit pricing and model-serving infrastructure answer different questions. Estimate GPU fit, rent, paid idle capacity, and API crossover for a specific model and traffic shape.
Can't find what you're looking for? Get in touch
How does the balance work?
You add credits with one-time top-ups, starting at $10, and spend from one balance, in dollars. Optional auto-recharge adds credits automatically when your balance runs low. Top-ups stay valid for one year, and the same balance covers agent plans, optimization, benchmarking, deploys, and hosted inference.
Describe the goal. RunInfra builds and optimizes the stack.