Consider FP8 or INT8 if the checkpoint and hardware support it
FP16 weights use 16.1 GB and bring total memory to 19.7 GB; FP8 or INT8 would lower weights to 8 GB and total memory to 9.1 GB, which fits inside 18 GB.
Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.
Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.
Model
meta-llama/Llama-3.1-8B-Instruct
GPU
NVIDIA RTX 4000 Ada 20 GB
Requests/day
50,000
GPU count
1
Quantization
FP16
Tokens/request
500 in, 500 out
Does not fit
The model and KV cache need more VRAM than the selected GPU count provides. Cost is intentionally hidden for a configuration that cannot run.
Model
meta-llama/Llama-3.1-8B-Instruct
Hardware
1 x NVIDIA RTX 4000 Ada 20 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
Memory-fit rental floors
Each row uses an engine GPU fit when present; the remaining catalog GPUs apply this result's required memory to published VRAM. GPU count rises only enough to clear memory.
Engine memory-fit result
Memory fit: 1x NVIDIA A10 24 GB at 90% vLLM GPU memory utilization.
On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The rentable GPU catalog is ranked by one-deployment monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. These are arithmetic floors, not provider offers. Throughput, replica demand, traffic capacity, topology support, and current availability are not established.
19.7 GB required, 21.6 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
One-deployment monthly rent floor
$116.88
RunPod, Community Cloud, on demand
Captured 2026-08-08
Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.
Open rental source19.7 GB required, 28.8 GB available inside the engine memory budget.
Current result memory total, Ampere, 448 GB/s memory bandwidth
One-deployment monthly rent floor
$219.15
Hyperstack, on demand
Captured 2026-08-08
Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.
Open rental source19.7 GB required, 43.2 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
One-deployment monthly rent floor
$241.07
RunPod, Community Cloud, on demand
Captured 2026-08-08
Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.
Open rental sourceComputed actions and cost signals appear only when the result supports them.
FP16 weights use 16.1 GB and bring total memory to 19.7 GB; FP8 or INT8 would lower weights to 8 GB and total memory to 9.1 GB, which fits inside 18 GB.
Weights use 16.1 GB and fit inside 18 GB, but the 1.1 GB KV cache raises total memory to 19.7 GB; reduce --max-model-len or peak concurrency before adding hardware.
The current vLLM command uses tensor parallel size 1 and provides 18 GB, while 19.7 GB requires 2 same-GPU workers for 36 GB; raise --tensor-parallel-size to 2 if the model topology supports it.
A physics bound, not an expected throughput result.
22 output tok/sCeiling
Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.
Decode ceiling derived from GPU memory bandwidth and model weight bytes, as of 2026-08-08
The exact command attached to this result.
vllm serve 'meta-llama/Llama-3.1-8B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9Every published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine feasibility runtime-overhead policy
As of 2026-08-09
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Does not fit
The model and KV cache need more VRAM than the selected GPU count provides. Cost is intentionally hidden for a configuration that cannot run.
Model
meta-llama/Llama-3.1-8B-Instruct
Hardware
1 x NVIDIA RTX 4000 Ada 20 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
Memory-fit rental floors
Each row uses an engine GPU fit when present; the remaining catalog GPUs apply this result's required memory to published VRAM. GPU count rises only enough to clear memory.
Engine memory-fit result
Memory fit: 1x NVIDIA A10 24 GB at 90% vLLM GPU memory utilization.
On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The rentable GPU catalog is ranked by one-deployment monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. These are arithmetic floors, not provider offers. Throughput, replica demand, traffic capacity, topology support, and current availability are not established.
19.7 GB required, 21.6 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
One-deployment monthly rent floor
$116.88
RunPod, Community Cloud, on demand
Captured 2026-08-08
Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.
Open rental source19.7 GB required, 28.8 GB available inside the engine memory budget.
Current result memory total, Ampere, 448 GB/s memory bandwidth
One-deployment monthly rent floor
$219.15
Hyperstack, on demand
Captured 2026-08-08
Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.
Open rental source19.7 GB required, 43.2 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
One-deployment monthly rent floor
$241.07
RunPod, Community Cloud, on demand
Captured 2026-08-08
Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.
Open rental sourceComputed actions and cost signals appear only when the result supports them.
FP16 weights use 16.1 GB and bring total memory to 19.7 GB; FP8 or INT8 would lower weights to 8 GB and total memory to 9.1 GB, which fits inside 18 GB.
Weights use 16.1 GB and fit inside 18 GB, but the 1.1 GB KV cache raises total memory to 19.7 GB; reduce --max-model-len or peak concurrency before adding hardware.
The current vLLM command uses tensor parallel size 1 and provides 18 GB, while 19.7 GB requires 2 same-GPU workers for 36 GB; raise --tensor-parallel-size to 2 if the model topology supports it.
A physics bound, not an expected throughput result.
22 output tok/sCeiling
Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.
Decode ceiling derived from GPU memory bandwidth and model weight bytes, as of 2026-08-08
The exact command attached to this result.
vllm serve 'meta-llama/Llama-3.1-8B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9Every published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine feasibility runtime-overhead policy
As of 2026-08-09
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Does not fit
The model and KV cache need more VRAM than the selected GPU count provides. Cost is intentionally hidden for a configuration that cannot run.
Model
meta-llama/Llama-3.1-8B-Instruct
Hardware
1 x NVIDIA RTX 4000 Ada 20 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
Memory-fit rental floors
Each row uses an engine GPU fit when present; the remaining catalog GPUs apply this result's required memory to published VRAM. GPU count rises only enough to clear memory.
Engine memory-fit result
Memory fit: 1x NVIDIA A10 24 GB at 90% vLLM GPU memory utilization.
On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The rentable GPU catalog is ranked by one-deployment monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. These are arithmetic floors, not provider offers. Throughput, replica demand, traffic capacity, topology support, and current availability are not established.
19.7 GB required, 21.6 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
One-deployment monthly rent floor
$116.88
RunPod, Community Cloud, on demand
Captured 2026-08-08
Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.
Open rental source19.7 GB required, 28.8 GB available inside the engine memory budget.
Current result memory total, Ampere, 448 GB/s memory bandwidth
One-deployment monthly rent floor
$219.15
Hyperstack, on demand
Captured 2026-08-08
Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.
Open rental source19.7 GB required, 43.2 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
One-deployment monthly rent floor
$241.07
RunPod, Community Cloud, on demand
Captured 2026-08-08
Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.
Open rental sourceComputed actions and cost signals appear only when the result supports them.
FP16 weights use 16.1 GB and bring total memory to 19.7 GB; FP8 or INT8 would lower weights to 8 GB and total memory to 9.1 GB, which fits inside 18 GB.
Weights use 16.1 GB and fit inside 18 GB, but the 1.1 GB KV cache raises total memory to 19.7 GB; reduce --max-model-len or peak concurrency before adding hardware.
The current vLLM command uses tensor parallel size 1 and provides 18 GB, while 19.7 GB requires 2 same-GPU workers for 36 GB; raise --tensor-parallel-size to 2 if the model topology supports it.
A physics bound, not an expected throughput result.
22 output tok/sCeiling
Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.
Decode ceiling derived from GPU memory bandwidth and model weight bytes, as of 2026-08-08
The exact command attached to this result.
vllm serve 'meta-llama/Llama-3.1-8B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9Every published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine feasibility runtime-overhead policy
As of 2026-08-09
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Does not fit
The model and KV cache need more VRAM than the selected GPU count provides. Cost is intentionally hidden for a configuration that cannot run.
Model
meta-llama/Llama-3.1-8B-Instruct
Hardware
1 x NVIDIA RTX 4000 Ada 20 GB, FP16
Weights and KV cache stay separate so the fit is reproducible.
Memory-fit rental floors
Each row uses an engine GPU fit when present; the remaining catalog GPUs apply this result's required memory to published VRAM. GPU count rises only enough to clear memory.
Engine memory-fit result
Memory fit: 1x NVIDIA A10 24 GB at 90% vLLM GPU memory utilization.
On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The rentable GPU catalog is ranked by one-deployment monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. These are arithmetic floors, not provider offers. Throughput, replica demand, traffic capacity, topology support, and current availability are not established.
19.7 GB required, 21.6 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
One-deployment monthly rent floor
$116.88
RunPod, Community Cloud, on demand
Captured 2026-08-08
Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.
Open rental source19.7 GB required, 28.8 GB available inside the engine memory budget.
Current result memory total, Ampere, 448 GB/s memory bandwidth
One-deployment monthly rent floor
$219.15
Hyperstack, on demand
Captured 2026-08-08
Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.
Open rental source19.7 GB required, 43.2 GB available inside the engine memory budget.
Current result memory total, Ampere, 768 GB/s memory bandwidth
One-deployment monthly rent floor
$241.07
RunPod, Community Cloud, on demand
Captured 2026-08-08
Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.
Open rental sourceComputed actions and cost signals appear only when the result supports them.
FP16 weights use 16.1 GB and bring total memory to 19.7 GB; FP8 or INT8 would lower weights to 8 GB and total memory to 9.1 GB, which fits inside 18 GB.
Weights use 16.1 GB and fit inside 18 GB, but the 1.1 GB KV cache raises total memory to 19.7 GB; reduce --max-model-len or peak concurrency before adding hardware.
The current vLLM command uses tensor parallel size 1 and provides 18 GB, while 19.7 GB requires 2 same-GPU workers for 36 GB; raise --tensor-parallel-size to 2 if the model topology supports it.
A physics bound, not an expected throughput result.
22 output tok/sCeiling
Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.
Decode ceiling derived from GPU memory bandwidth and model weight bytes, as of 2026-08-08
The exact command attached to this result.
vllm serve 'meta-llama/Llama-3.1-8B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9Every published number keeps its source and date.
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
RunInfra Engine feasibility runtime-overhead policy
As of 2026-08-09
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
As of 2026-08-08
Embed code appears only when the calculator can publish a sourced cost for a fitting configuration.
Badge unavailable for this configuration
This configuration does not fit the selected hardware, so the calculator does not publish a cost or embed code.
© 2026 RunInfra. All rights reserved.