Loading the fit, utilization assumptions, and cited price data.
Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.
Model
Qwen/Qwen2.5-14B-Instruct
GPU
NVIDIA H100 SXM 80 GB
Requests/day
50,000
GPU count
1
Quantization
FP16
Tokens/request
500 in, 500 out
Priced configuration
No comparable dated API rate is available for this model. The self-host bill remains reproducible without inventing an API baseline.
Model
Qwen/Qwen2.5-14B-Instruct
Hardware
1 x NVIDIA H100 SXM 80 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
The exact command attached to this result.
vllm serve 'Qwen/Qwen2.5-14B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 32768 --gpu-memory-utilization 0.9Utilization and idle share are shown before the token price.
Paid idle capacity
14.7%
321.3 of 2,191.5 billed GPU-hours sit idle. The provider still charges every hour.
Busy GPU-hours
1,870.2
Useful work and paid idle time sum to the billed total above.
CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08
At full saturation
$6.59 / 1M output tokens
At actual utilization
$7.75 / 1M output tokens
Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.
Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.
Scroll horizontally to read every volume point.
500 requests per day, 7,609,375 output tokens per month: self-host $1,965.05, no comparable API bill. 1,500 requests per day, 22,828,125 output tokens per month: self-host $1,965.05, no comparable API bill. 5,000 requests per day, 76,093,750 output tokens per month: self-host $1,965.05, no comparable API bill. 15,000 requests per day, 228,281,250 output tokens per month: self-host $1,965.05, no comparable API bill. 19,335 requests per day, 294,254,193 output tokens per month: self-host $1,965.05, no comparable API bill. 19,726 requests per day, 300,198,722 output tokens per month: self-host $3,930.09, no comparable API bill. 50,000 requests per day, 760,937,500 output tokens per month: self-host $5,895.14, no comparable API bill. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $15,720.36, no comparable API bill. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $51,091.17, no comparable API bill. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $151,308.47, no comparable API bill. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $505,016.57, no comparable API bill.
Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.
2,191.5 billed GPU-hours, split by the engine into useful work and paid idle capacity.
Scroll horizontally to read every chart label and value.
Busy GPU-hours: 1,870.2 hrs, Engine-reported share doing token work. Idle GPU-hours billed: 321.3 hrs, Engine-reported share paid while idle.
Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.
USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.
Scroll horizontally to read every chart label and value.
DigitalOcean, GPU Droplets: $4.41/hr, on demand GPU rate, captured 2026-08-08. Hyperstack: $3.20/hr, on demand GPU rate, captured 2026-08-08. Lambda Cloud: $4.29/hr, on demand GPU rate, captured 2026-08-08. Nebius: $3.85/hr, on demand GPU rate, captured 2026-08-08. Cheapest + Used: RunPod, Community Cloud: $2.69/hr, on demand GPU rate, captured 2026-08-08, used by the engine for this calculation. RunPod, Secure Cloud: $2.99/hr, on demand GPU rate, captured 2026-08-08. Verda: $3.25/hr, on demand GPU rate, captured 2026-08-08.
Every rate is for the selected GPU and keeps its provider, tier, source, and capture date. A price without a date is not included.
No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine parameter-only KV-cache heuristic
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Priced configuration
No comparable dated API rate is available for this model. The self-host bill remains reproducible without inventing an API baseline.
Model
Qwen/Qwen2.5-14B-Instruct
Hardware
1 x NVIDIA H100 SXM 80 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
The exact command attached to this result.
vllm serve 'Qwen/Qwen2.5-14B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 32768 --gpu-memory-utilization 0.9Utilization and idle share are shown before the token price.
Paid idle capacity
14.7%
321.3 of 2,191.5 billed GPU-hours sit idle. The provider still charges every hour.
Busy GPU-hours
1,870.2
Useful work and paid idle time sum to the billed total above.
CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08
At full saturation
$6.59 / 1M output tokens
At actual utilization
$7.75 / 1M output tokens
Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.
Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.
Scroll horizontally to read every volume point.
500 requests per day, 7,609,375 output tokens per month: self-host $1,965.05, no comparable API bill. 1,500 requests per day, 22,828,125 output tokens per month: self-host $1,965.05, no comparable API bill. 5,000 requests per day, 76,093,750 output tokens per month: self-host $1,965.05, no comparable API bill. 15,000 requests per day, 228,281,250 output tokens per month: self-host $1,965.05, no comparable API bill. 19,335 requests per day, 294,254,193 output tokens per month: self-host $1,965.05, no comparable API bill. 19,726 requests per day, 300,198,722 output tokens per month: self-host $3,930.09, no comparable API bill. 50,000 requests per day, 760,937,500 output tokens per month: self-host $5,895.14, no comparable API bill. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $15,720.36, no comparable API bill. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $51,091.17, no comparable API bill. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $151,308.47, no comparable API bill. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $505,016.57, no comparable API bill.
Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.
2,191.5 billed GPU-hours, split by the engine into useful work and paid idle capacity.
Scroll horizontally to read every chart label and value.
Busy GPU-hours: 1,870.2 hrs, Engine-reported share doing token work. Idle GPU-hours billed: 321.3 hrs, Engine-reported share paid while idle.
Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.
USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.
Scroll horizontally to read every chart label and value.
DigitalOcean, GPU Droplets: $4.41/hr, on demand GPU rate, captured 2026-08-08. Hyperstack: $3.20/hr, on demand GPU rate, captured 2026-08-08. Lambda Cloud: $4.29/hr, on demand GPU rate, captured 2026-08-08. Nebius: $3.85/hr, on demand GPU rate, captured 2026-08-08. Cheapest + Used: RunPod, Community Cloud: $2.69/hr, on demand GPU rate, captured 2026-08-08, used by the engine for this calculation. RunPod, Secure Cloud: $2.99/hr, on demand GPU rate, captured 2026-08-08. Verda: $3.25/hr, on demand GPU rate, captured 2026-08-08.
Every rate is for the selected GPU and keeps its provider, tier, source, and capture date. A price without a date is not included.
No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine parameter-only KV-cache heuristic
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Priced configuration
No comparable dated API rate is available for this model. The self-host bill remains reproducible without inventing an API baseline.
Model
Qwen/Qwen2.5-14B-Instruct
Hardware
1 x NVIDIA H100 SXM 80 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
The exact command attached to this result.
vllm serve 'Qwen/Qwen2.5-14B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 32768 --gpu-memory-utilization 0.9Utilization and idle share are shown before the token price.
Paid idle capacity
14.7%
321.3 of 2,191.5 billed GPU-hours sit idle. The provider still charges every hour.
Busy GPU-hours
1,870.2
Useful work and paid idle time sum to the billed total above.
CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08
At full saturation
$6.59 / 1M output tokens
At actual utilization
$7.75 / 1M output tokens
Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.
Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.
Scroll horizontally to read every volume point.
500 requests per day, 7,609,375 output tokens per month: self-host $1,965.05, no comparable API bill. 1,500 requests per day, 22,828,125 output tokens per month: self-host $1,965.05, no comparable API bill. 5,000 requests per day, 76,093,750 output tokens per month: self-host $1,965.05, no comparable API bill. 15,000 requests per day, 228,281,250 output tokens per month: self-host $1,965.05, no comparable API bill. 19,335 requests per day, 294,254,193 output tokens per month: self-host $1,965.05, no comparable API bill. 19,726 requests per day, 300,198,722 output tokens per month: self-host $3,930.09, no comparable API bill. 50,000 requests per day, 760,937,500 output tokens per month: self-host $5,895.14, no comparable API bill. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $15,720.36, no comparable API bill. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $51,091.17, no comparable API bill. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $151,308.47, no comparable API bill. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $505,016.57, no comparable API bill.
Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.
2,191.5 billed GPU-hours, split by the engine into useful work and paid idle capacity.
Scroll horizontally to read every chart label and value.
Busy GPU-hours: 1,870.2 hrs, Engine-reported share doing token work. Idle GPU-hours billed: 321.3 hrs, Engine-reported share paid while idle.
Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.
USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.
Scroll horizontally to read every chart label and value.
DigitalOcean, GPU Droplets: $4.41/hr, on demand GPU rate, captured 2026-08-08. Hyperstack: $3.20/hr, on demand GPU rate, captured 2026-08-08. Lambda Cloud: $4.29/hr, on demand GPU rate, captured 2026-08-08. Nebius: $3.85/hr, on demand GPU rate, captured 2026-08-08. Cheapest + Used: RunPod, Community Cloud: $2.69/hr, on demand GPU rate, captured 2026-08-08, used by the engine for this calculation. RunPod, Secure Cloud: $2.99/hr, on demand GPU rate, captured 2026-08-08. Verda: $3.25/hr, on demand GPU rate, captured 2026-08-08.
Every rate is for the selected GPU and keeps its provider, tier, source, and capture date. A price without a date is not included.
No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine parameter-only KV-cache heuristic
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Priced configuration
No comparable dated API rate is available for this model. The self-host bill remains reproducible without inventing an API baseline.
Model
Qwen/Qwen2.5-14B-Instruct
Hardware
1 x NVIDIA H100 SXM 80 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
The exact command attached to this result.
vllm serve 'Qwen/Qwen2.5-14B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 32768 --gpu-memory-utilization 0.9Utilization and idle share are shown before the token price.
Paid idle capacity
14.7%
321.3 of 2,191.5 billed GPU-hours sit idle. The provider still charges every hour.
Busy GPU-hours
1,870.2
Useful work and paid idle time sum to the billed total above.
CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08
At full saturation
$6.59 / 1M output tokens
At actual utilization
$7.75 / 1M output tokens
Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.
Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.
Scroll horizontally to read every volume point.
500 requests per day, 7,609,375 output tokens per month: self-host $1,965.05, no comparable API bill. 1,500 requests per day, 22,828,125 output tokens per month: self-host $1,965.05, no comparable API bill. 5,000 requests per day, 76,093,750 output tokens per month: self-host $1,965.05, no comparable API bill. 15,000 requests per day, 228,281,250 output tokens per month: self-host $1,965.05, no comparable API bill. 19,335 requests per day, 294,254,193 output tokens per month: self-host $1,965.05, no comparable API bill. 19,726 requests per day, 300,198,722 output tokens per month: self-host $3,930.09, no comparable API bill. 50,000 requests per day, 760,937,500 output tokens per month: self-host $5,895.14, no comparable API bill. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $15,720.36, no comparable API bill. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $51,091.17, no comparable API bill. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $151,308.47, no comparable API bill. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $505,016.57, no comparable API bill.
Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.
2,191.5 billed GPU-hours, split by the engine into useful work and paid idle capacity.
Scroll horizontally to read every chart label and value.
Busy GPU-hours: 1,870.2 hrs, Engine-reported share doing token work. Idle GPU-hours billed: 321.3 hrs, Engine-reported share paid while idle.
Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.
USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.
Scroll horizontally to read every chart label and value.
DigitalOcean, GPU Droplets: $4.41/hr, on demand GPU rate, captured 2026-08-08. Hyperstack: $3.20/hr, on demand GPU rate, captured 2026-08-08. Lambda Cloud: $4.29/hr, on demand GPU rate, captured 2026-08-08. Nebius: $3.85/hr, on demand GPU rate, captured 2026-08-08. Cheapest + Used: RunPod, Community Cloud: $2.69/hr, on demand GPU rate, captured 2026-08-08, used by the engine for this calculation. RunPod, Secure Cloud: $2.99/hr, on demand GPU rate, captured 2026-08-08. Verda: $3.25/hr, on demand GPU rate, captured 2026-08-08.
Every rate is for the selected GPU and keeps its provider, tier, source, and capture date. A price without a date is not included.
No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine parameter-only KV-cache heuristic
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
The badge keeps its evidence label and links to the full result, sources, and caveats.
HTML
<a href="https://runinfra.ai/calc/qwen-2.5-14b/h100/50k-req-day"><img src="https://runinfra.ai/api/calc/badge?hfId=Qwen%2FQwen2.5-14B-Instruct&gpuId=H100&gpuCount=1&requestsPerDay=50000&avgInputTokens=500&avgOutputTokens=500&utilization=0.3&quantization=fp16" alt="Calculated by RunInfra"></a>Markdown
[](https://runinfra.ai/calc/qwen-2.5-14b/h100/50k-req-day)© 2026 RunInfra. All rights reserved.