Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.
Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.
Model
microsoft/Phi-3.5-mini-instruct
GPU
NVIDIA H100 NVL 94 GB
Requests/day
50,000
GPU count
1
Quantization
FP16
Tokens/request
500 in, 500 out
Estimated memory fit
The estimated fit verdict, compatible GPUs, and reproducing command are still useful. Money stays absent until this traffic shape has a complete measured, cited, or physics-bound evidence path.
Model
microsoft/Phi-3.5-mini-instruct
Hardware
1 x NVIDIA H100 NVL 94 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
A physics bound, not an expected throughput result.
510 output tok/sCeiling
Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.
Decode ceiling derived from GPU memory bandwidth and model weight bytes, as of 2026-08-08
Dated GPU-hour rental prices for this hardware. They are not monthly serving totals because throughput evidence is missing.
Scroll horizontally to read every chart label and value.
Microsoft Azure, East US 2: $6.98, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest: RunPod, Community Cloud: $2.59, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). RunPod, Secure Cloud: $3.19, on demand GPU-hour rate, captured 2026-08-08 (1 day ago).
These are GPU-hour rental prices only. A monthly serving total requires throughput evidence and remains absent.
The exact command attached to this result.
vllm serve 'microsoft/Phi-3.5-mini-instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine feasibility runtime-overhead policy
As of 2026-08-09
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Estimated memory fit
The estimated fit verdict, compatible GPUs, and reproducing command are still useful. Money stays absent until this traffic shape has a complete measured, cited, or physics-bound evidence path.
Model
microsoft/Phi-3.5-mini-instruct
Hardware
1 x NVIDIA H100 NVL 94 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
A physics bound, not an expected throughput result.
510 output tok/sCeiling
Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.
Decode ceiling derived from GPU memory bandwidth and model weight bytes, as of 2026-08-08
Dated GPU-hour rental prices for this hardware. They are not monthly serving totals because throughput evidence is missing.
Scroll horizontally to read every chart label and value.
Microsoft Azure, East US 2: $6.98, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest: RunPod, Community Cloud: $2.59, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). RunPod, Secure Cloud: $3.19, on demand GPU-hour rate, captured 2026-08-08 (1 day ago).
These are GPU-hour rental prices only. A monthly serving total requires throughput evidence and remains absent.
The exact command attached to this result.
vllm serve 'microsoft/Phi-3.5-mini-instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine feasibility runtime-overhead policy
As of 2026-08-09
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Estimated memory fit
The estimated fit verdict, compatible GPUs, and reproducing command are still useful. Money stays absent until this traffic shape has a complete measured, cited, or physics-bound evidence path.
Model
microsoft/Phi-3.5-mini-instruct
Hardware
1 x NVIDIA H100 NVL 94 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
A physics bound, not an expected throughput result.
510 output tok/sCeiling
Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.
Decode ceiling derived from GPU memory bandwidth and model weight bytes, as of 2026-08-08
Dated GPU-hour rental prices for this hardware. They are not monthly serving totals because throughput evidence is missing.
Scroll horizontally to read every chart label and value.
Microsoft Azure, East US 2: $6.98, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest: RunPod, Community Cloud: $2.59, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). RunPod, Secure Cloud: $3.19, on demand GPU-hour rate, captured 2026-08-08 (1 day ago).
These are GPU-hour rental prices only. A monthly serving total requires throughput evidence and remains absent.
The exact command attached to this result.
vllm serve 'microsoft/Phi-3.5-mini-instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine feasibility runtime-overhead policy
As of 2026-08-09
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Estimated memory fit
The estimated fit verdict, compatible GPUs, and reproducing command are still useful. Money stays absent until this traffic shape has a complete measured, cited, or physics-bound evidence path.
Model
microsoft/Phi-3.5-mini-instruct
Hardware
1 x NVIDIA H100 NVL 94 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
A physics bound, not an expected throughput result.
510 output tok/sCeiling
Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.
Decode ceiling derived from GPU memory bandwidth and model weight bytes, as of 2026-08-08
Dated GPU-hour rental prices for this hardware. They are not monthly serving totals because throughput evidence is missing.
Scroll horizontally to read every chart label and value.
Microsoft Azure, East US 2: $6.98, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest: RunPod, Community Cloud: $2.59, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). RunPod, Secure Cloud: $3.19, on demand GPU-hour rate, captured 2026-08-08 (1 day ago).
These are GPU-hour rental prices only. A monthly serving total requires throughput evidence and remains absent.
The exact command attached to this result.
vllm serve 'microsoft/Phi-3.5-mini-instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.
Run this configurationEvery published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine feasibility runtime-overhead policy
As of 2026-08-09
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Embed code appears only when the calculator can publish a sourced cost for a fitting configuration.
Badge unavailable for this configuration
The calculator does not have a complete throughput source for this workload, so it withheld the cost and does not generate embed code.
© 2026 RunInfra. All rights reserved.