Move this workload off oversized dedicated capacity
91.4% of 730.5 billed GPU-hours are idle at 8.6% derived utilization; choose a smaller fitting GPU or move to shared or serverless capacity rather than adding hardware.
Loading the fit, utilization assumptions, and cited price data.
Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.
Model
Qwen/Qwen2.5-0.5B-Instruct
GPU
NVIDIA H100 SXM 80 GB
Requests/day
50,000
GPU count
1
Quantization
FP16
Tokens/request
500 in, 500 out
Priced configuration
No comparable dated API rate is available for this model. The self-host bill remains reproducible without inventing an API baseline.
Model
Qwen/Qwen2.5-0.5B-Instruct
Hardware
1 x NVIDIA H100 SXM 80 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
The exact command attached to this result.
vllm serve 'Qwen/Qwen2.5-0.5B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 32768 --gpu-memory-utilization 0.9Utilization and idle share are shown before the token price.
Paid idle capacity
91.4%
667.9 of 730.5 billed GPU-hours sit idle. The provider still charges every hour.
Busy GPU-hours
62.6
Useful work and paid idle time sum to the billed total above.
CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08
At full saturation
$0.22 / 1M output tokens
At actual utilization
$2.58 / 1M output tokens
Hardware recommendation
91.4% of billed GPU-hours are idle, so the shortlist keeps only smaller memory-fitting rentals with a lower monthly amount at the current replica count.
On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The complete rentable GPU catalog is ranked only by monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. The engine's current 1 replica is held constant; alternative throughput and replica demand are not re-estimated. Memory fit does not establish tensor-parallel topology support.
1.4 GB required, 14.4 GB available inside the engine memory budget.
Current result memory total, Ampere, 448 GB/s memory bandwidth
Monthly rent at current replica count
$109.58
Hyperstack, on demand
Captured 2026-08-08
At the selected 10% utilization scenario, the provider still bills the full monthly amount.
Open rental source1.4 GB required, 21.6 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
Monthly rent at current replica count
$116.88
RunPod, Community Cloud, on demand
Captured 2026-08-08
At the selected 10% utilization scenario, the provider still bills the full monthly amount.
Open rental source1.4 GB required, 43.2 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
Monthly rent at current replica count
$241.07
RunPod, Community Cloud, on demand
Captured 2026-08-08
At the selected 10% utilization scenario, the provider still bills the full monthly amount.
Open rental sourceOnly result-triggered recommendations appear here.
91.4% of 730.5 billed GPU-hours are idle at 8.6% derived utilization; choose a smaller fitting GPU or move to shared or serverless capacity rather than adding hardware.
Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.
Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.
Scroll horizontally to read every volume point.
500 requests per day, 7,609,375 output tokens per month: self-host $1,965.05, no comparable API bill. 1,500 requests per day, 22,828,125 output tokens per month: self-host $1,965.05, no comparable API bill. 5,000 requests per day, 76,093,750 output tokens per month: self-host $1,965.05, no comparable API bill. 15,000 requests per day, 228,281,250 output tokens per month: self-host $1,965.05, no comparable API bill. 50,000 requests per day, 760,937,500 output tokens per month: self-host $1,965.05, no comparable API bill. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $1,965.05, no comparable API bill. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $1,965.05, no comparable API bill. 578,055 requests per day, 8,797,279,475 output tokens per month: self-host $1,965.05, no comparable API bill. 589,733 requests per day, 8,975,002,292 output tokens per month: self-host $3,930.09, no comparable API bill. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $5,895.14, no comparable API bill. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $17,685.41, no comparable API bill.
Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.
730.5 billed GPU-hours, split by the engine into useful work and paid idle capacity.
Scroll horizontally to read every chart label and value.
Busy GPU-hours: 62.6 hrs, Engine-reported share doing token work. Idle GPU-hours billed: 667.9 hrs, Engine-reported share paid while idle.
Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.
USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.
Scroll horizontally to read every chart label and value.
DigitalOcean, GPU Droplets: $4.41/hr, on demand GPU rate, captured 2026-08-08. Hyperstack: $3.20/hr, on demand GPU rate, captured 2026-08-08. Lambda Cloud: $4.29/hr, on demand GPU rate, captured 2026-08-08. Nebius: $3.85/hr, on demand GPU rate, captured 2026-08-08. Cheapest + Used: RunPod, Community Cloud: $2.69/hr, on demand GPU rate, captured 2026-08-08, used by the engine for this calculation. RunPod, Secure Cloud: $2.99/hr, on demand GPU rate, captured 2026-08-08. Verda: $3.25/hr, on demand GPU rate, captured 2026-08-08.
Every rate is for the selected GPU and keeps its provider, tier, source, and capture date. A price without a date is not included.
No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Priced configuration
No comparable dated API rate is available for this model. The self-host bill remains reproducible without inventing an API baseline.
Model
Qwen/Qwen2.5-0.5B-Instruct
Hardware
1 x NVIDIA H100 SXM 80 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
The exact command attached to this result.
vllm serve 'Qwen/Qwen2.5-0.5B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 32768 --gpu-memory-utilization 0.9Utilization and idle share are shown before the token price.
Paid idle capacity
91.4%
667.9 of 730.5 billed GPU-hours sit idle. The provider still charges every hour.
Busy GPU-hours
62.6
Useful work and paid idle time sum to the billed total above.
CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08
At full saturation
$0.22 / 1M output tokens
At actual utilization
$2.58 / 1M output tokens
Hardware recommendation
91.4% of billed GPU-hours are idle, so the shortlist keeps only smaller memory-fitting rentals with a lower monthly amount at the current replica count.
On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The complete rentable GPU catalog is ranked only by monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. The engine's current 1 replica is held constant; alternative throughput and replica demand are not re-estimated. Memory fit does not establish tensor-parallel topology support.
1.4 GB required, 14.4 GB available inside the engine memory budget.
Current result memory total, Ampere, 448 GB/s memory bandwidth
Monthly rent at current replica count
$109.58
Hyperstack, on demand
Captured 2026-08-08
At the selected 30% utilization scenario, the provider still bills the full monthly amount.
Open rental source1.4 GB required, 21.6 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
Monthly rent at current replica count
$116.88
RunPod, Community Cloud, on demand
Captured 2026-08-08
At the selected 30% utilization scenario, the provider still bills the full monthly amount.
Open rental source1.4 GB required, 43.2 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
Monthly rent at current replica count
$241.07
RunPod, Community Cloud, on demand
Captured 2026-08-08
At the selected 30% utilization scenario, the provider still bills the full monthly amount.
Open rental sourceOnly result-triggered recommendations appear here.
91.4% of 730.5 billed GPU-hours are idle at 8.6% derived utilization; choose a smaller fitting GPU or move to shared or serverless capacity rather than adding hardware.
Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.
Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.
Scroll horizontally to read every volume point.
500 requests per day, 7,609,375 output tokens per month: self-host $1,965.05, no comparable API bill. 1,500 requests per day, 22,828,125 output tokens per month: self-host $1,965.05, no comparable API bill. 5,000 requests per day, 76,093,750 output tokens per month: self-host $1,965.05, no comparable API bill. 15,000 requests per day, 228,281,250 output tokens per month: self-host $1,965.05, no comparable API bill. 50,000 requests per day, 760,937,500 output tokens per month: self-host $1,965.05, no comparable API bill. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $1,965.05, no comparable API bill. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $1,965.05, no comparable API bill. 578,055 requests per day, 8,797,279,475 output tokens per month: self-host $1,965.05, no comparable API bill. 589,733 requests per day, 8,975,002,292 output tokens per month: self-host $3,930.09, no comparable API bill. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $5,895.14, no comparable API bill. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $17,685.41, no comparable API bill.
Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.
730.5 billed GPU-hours, split by the engine into useful work and paid idle capacity.
Scroll horizontally to read every chart label and value.
Busy GPU-hours: 62.6 hrs, Engine-reported share doing token work. Idle GPU-hours billed: 667.9 hrs, Engine-reported share paid while idle.
Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.
USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.
Scroll horizontally to read every chart label and value.
DigitalOcean, GPU Droplets: $4.41/hr, on demand GPU rate, captured 2026-08-08. Hyperstack: $3.20/hr, on demand GPU rate, captured 2026-08-08. Lambda Cloud: $4.29/hr, on demand GPU rate, captured 2026-08-08. Nebius: $3.85/hr, on demand GPU rate, captured 2026-08-08. Cheapest + Used: RunPod, Community Cloud: $2.69/hr, on demand GPU rate, captured 2026-08-08, used by the engine for this calculation. RunPod, Secure Cloud: $2.99/hr, on demand GPU rate, captured 2026-08-08. Verda: $3.25/hr, on demand GPU rate, captured 2026-08-08.
Every rate is for the selected GPU and keeps its provider, tier, source, and capture date. A price without a date is not included.
No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Priced configuration
No comparable dated API rate is available for this model. The self-host bill remains reproducible without inventing an API baseline.
Model
Qwen/Qwen2.5-0.5B-Instruct
Hardware
1 x NVIDIA H100 SXM 80 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
The exact command attached to this result.
vllm serve 'Qwen/Qwen2.5-0.5B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 32768 --gpu-memory-utilization 0.9Utilization and idle share are shown before the token price.
Paid idle capacity
91.4%
667.9 of 730.5 billed GPU-hours sit idle. The provider still charges every hour.
Busy GPU-hours
62.6
Useful work and paid idle time sum to the billed total above.
CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08
At full saturation
$0.22 / 1M output tokens
At actual utilization
$2.58 / 1M output tokens
Hardware recommendation
91.4% of billed GPU-hours are idle, so the shortlist keeps only smaller memory-fitting rentals with a lower monthly amount at the current replica count.
On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The complete rentable GPU catalog is ranked only by monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. The engine's current 1 replica is held constant; alternative throughput and replica demand are not re-estimated. Memory fit does not establish tensor-parallel topology support.
1.4 GB required, 14.4 GB available inside the engine memory budget.
Current result memory total, Ampere, 448 GB/s memory bandwidth
Monthly rent at current replica count
$109.58
Hyperstack, on demand
Captured 2026-08-08
At the selected 60% utilization scenario, the provider still bills the full monthly amount.
Open rental source1.4 GB required, 21.6 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
Monthly rent at current replica count
$116.88
RunPod, Community Cloud, on demand
Captured 2026-08-08
At the selected 60% utilization scenario, the provider still bills the full monthly amount.
Open rental source1.4 GB required, 43.2 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
Monthly rent at current replica count
$241.07
RunPod, Community Cloud, on demand
Captured 2026-08-08
At the selected 60% utilization scenario, the provider still bills the full monthly amount.
Open rental sourceOnly result-triggered recommendations appear here.
91.4% of 730.5 billed GPU-hours are idle at 8.6% derived utilization; choose a smaller fitting GPU or move to shared or serverless capacity rather than adding hardware.
Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.
Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.
Scroll horizontally to read every volume point.
500 requests per day, 7,609,375 output tokens per month: self-host $1,965.05, no comparable API bill. 1,500 requests per day, 22,828,125 output tokens per month: self-host $1,965.05, no comparable API bill. 5,000 requests per day, 76,093,750 output tokens per month: self-host $1,965.05, no comparable API bill. 15,000 requests per day, 228,281,250 output tokens per month: self-host $1,965.05, no comparable API bill. 50,000 requests per day, 760,937,500 output tokens per month: self-host $1,965.05, no comparable API bill. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $1,965.05, no comparable API bill. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $1,965.05, no comparable API bill. 578,055 requests per day, 8,797,279,475 output tokens per month: self-host $1,965.05, no comparable API bill. 589,733 requests per day, 8,975,002,292 output tokens per month: self-host $3,930.09, no comparable API bill. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $5,895.14, no comparable API bill. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $17,685.41, no comparable API bill.
Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.
730.5 billed GPU-hours, split by the engine into useful work and paid idle capacity.
Scroll horizontally to read every chart label and value.
Busy GPU-hours: 62.6 hrs, Engine-reported share doing token work. Idle GPU-hours billed: 667.9 hrs, Engine-reported share paid while idle.
Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.
USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.
Scroll horizontally to read every chart label and value.
DigitalOcean, GPU Droplets: $4.41/hr, on demand GPU rate, captured 2026-08-08. Hyperstack: $3.20/hr, on demand GPU rate, captured 2026-08-08. Lambda Cloud: $4.29/hr, on demand GPU rate, captured 2026-08-08. Nebius: $3.85/hr, on demand GPU rate, captured 2026-08-08. Cheapest + Used: RunPod, Community Cloud: $2.69/hr, on demand GPU rate, captured 2026-08-08, used by the engine for this calculation. RunPod, Secure Cloud: $2.99/hr, on demand GPU rate, captured 2026-08-08. Verda: $3.25/hr, on demand GPU rate, captured 2026-08-08.
Every rate is for the selected GPU and keeps its provider, tier, source, and capture date. A price without a date is not included.
No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Priced configuration
No comparable dated API rate is available for this model. The self-host bill remains reproducible without inventing an API baseline.
Model
Qwen/Qwen2.5-0.5B-Instruct
Hardware
1 x NVIDIA H100 SXM 80 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
The exact command attached to this result.
vllm serve 'Qwen/Qwen2.5-0.5B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 32768 --gpu-memory-utilization 0.9Utilization and idle share are shown before the token price.
Paid idle capacity
91.4%
667.9 of 730.5 billed GPU-hours sit idle. The provider still charges every hour.
Busy GPU-hours
62.6
Useful work and paid idle time sum to the billed total above.
CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08
At full saturation
$0.22 / 1M output tokens
At actual utilization
$2.58 / 1M output tokens
Hardware recommendation
91.4% of billed GPU-hours are idle, so the shortlist keeps only smaller memory-fitting rentals with a lower monthly amount at the current replica count.
On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The complete rentable GPU catalog is ranked only by monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. The engine's current 1 replica is held constant; alternative throughput and replica demand are not re-estimated. Memory fit does not establish tensor-parallel topology support.
1.4 GB required, 14.4 GB available inside the engine memory budget.
Current result memory total, Ampere, 448 GB/s memory bandwidth
Monthly rent at current replica count
$109.58
Hyperstack, on demand
Captured 2026-08-08
At the selected 90% utilization scenario, the provider still bills the full monthly amount.
Open rental source1.4 GB required, 21.6 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
Monthly rent at current replica count
$116.88
RunPod, Community Cloud, on demand
Captured 2026-08-08
At the selected 90% utilization scenario, the provider still bills the full monthly amount.
Open rental source1.4 GB required, 43.2 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
Monthly rent at current replica count
$241.07
RunPod, Community Cloud, on demand
Captured 2026-08-08
At the selected 90% utilization scenario, the provider still bills the full monthly amount.
Open rental sourceOnly result-triggered recommendations appear here.
91.4% of 730.5 billed GPU-hours are idle at 8.6% derived utilization; choose a smaller fitting GPU or move to shared or serverless capacity rather than adding hardware.
Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.
Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.
Scroll horizontally to read every volume point.
500 requests per day, 7,609,375 output tokens per month: self-host $1,965.05, no comparable API bill. 1,500 requests per day, 22,828,125 output tokens per month: self-host $1,965.05, no comparable API bill. 5,000 requests per day, 76,093,750 output tokens per month: self-host $1,965.05, no comparable API bill. 15,000 requests per day, 228,281,250 output tokens per month: self-host $1,965.05, no comparable API bill. 50,000 requests per day, 760,937,500 output tokens per month: self-host $1,965.05, no comparable API bill. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $1,965.05, no comparable API bill. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $1,965.05, no comparable API bill. 578,055 requests per day, 8,797,279,475 output tokens per month: self-host $1,965.05, no comparable API bill. 589,733 requests per day, 8,975,002,292 output tokens per month: self-host $3,930.09, no comparable API bill. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $5,895.14, no comparable API bill. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $17,685.41, no comparable API bill.
Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.
730.5 billed GPU-hours, split by the engine into useful work and paid idle capacity.
Scroll horizontally to read every chart label and value.
Busy GPU-hours: 62.6 hrs, Engine-reported share doing token work. Idle GPU-hours billed: 667.9 hrs, Engine-reported share paid while idle.
Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.
USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.
Scroll horizontally to read every chart label and value.
DigitalOcean, GPU Droplets: $4.41/hr, on demand GPU rate, captured 2026-08-08. Hyperstack: $3.20/hr, on demand GPU rate, captured 2026-08-08. Lambda Cloud: $4.29/hr, on demand GPU rate, captured 2026-08-08. Nebius: $3.85/hr, on demand GPU rate, captured 2026-08-08. Cheapest + Used: RunPod, Community Cloud: $2.69/hr, on demand GPU rate, captured 2026-08-08, used by the engine for this calculation. RunPod, Secure Cloud: $2.99/hr, on demand GPU rate, captured 2026-08-08. Verda: $3.25/hr, on demand GPU rate, captured 2026-08-08.
Every rate is for the selected GPU and keeps its provider, tier, source, and capture date. A price without a date is not included.
No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
The badge keeps its evidence label and links to the full result, sources, and caveats.
HTML
<a href="https://runinfra.ai/calc/qwen-2.5-0.5b/h100/50k-req-day"><img src="https://runinfra.ai/api/calc/badge?hfId=Qwen%2FQwen2.5-0.5B-Instruct&gpuId=H100&gpuCount=1&requestsPerDay=50000&avgInputTokens=500&avgOutputTokens=500&utilization=0.3&quantization=fp16" alt="Calculated by RunInfra"></a>Markdown
[](https://runinfra.ai/calc/qwen-2.5-0.5b/h100/50k-req-day)© 2026 RunInfra. All rights reserved.