Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.
Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.
Model
Qwen/Qwen2.5-14B-Instruct
GPU
NVIDIA A100 PCIe 80 GB
Requests/day
50,000
GPU count
1
Quantization
FP16
Tokens/request
500 in, 500 out
Lower-bound floor
The throughput basis is a 100%-utilization ceiling. Real throughput can be lower and real rent can be higher, so API comparisons and cost curves stay absent.
Model
Qwen/Qwen2.5-14B-Instruct
Hardware
1 x NVIDIA A100 PCIe 80 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
This line starts the selected vLLM shape. It is not a throughput reproduction claim.
vllm serve 'Qwen/Qwen2.5-14B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9A floor derived from a 100%-utilization throughput ceiling, not an expected bill.
At least per month
$4,346.47
Real throughput can be lower, which can require more replicas and raise rent.
Minimum replicas
5
Lower bound from ideal ceiling capacity. No utilization or idle-hour estimate is published.
Rental basis
$1.19/GPU-hour
RunPod, Community Cloud, on demand
Captured 2026-08-08 (1 day ago)
Ceiling capacity: 65 output tok/s at concurrency 1.
CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. A rental figure derived from this ceiling is a lower-bound floor, not an estimate; real throughput can be lower and real rent can be higher.
Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08
No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine feasibility runtime-overhead policy
As of 2026-08-09
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
The badge keeps its evidence label and links to the full result, sources, and caveats.
HTML
<a href="https://runinfra.ai/calc/qwen-2.5-14b/a100-80gb-pcie/50k-req-day"><img src="https://runinfra.ai/api/calc/badge?hfId=Qwen%2FQwen2.5-14B-Instruct&gpuId=A100-80GB-PCIe&gpuCount=1&requestsPerDay=50000&avgInputTokens=500&avgOutputTokens=500&utilization=0.3&quantization=fp16&servedContextTokens=8192" alt="Calculated by RunInfra"></a>Markdown
[](https://runinfra.ai/calc/qwen-2.5-14b/a100-80gb-pcie/50k-req-day)© 2026 RunInfra. All rights reserved.