RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

See what this model actually costs to serve.

Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.

The modelled monthly GPU rent for Llama-3.1-8B-Instruct using 8 replicas of 1 x NVIDIA A10 24 GB is $5,844.00 at 50,000 requests/day.

Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.

Model

meta-llama/Llama-3.1-8B-Instruct

GPU

NVIDIA A10 24 GB

Requests/day

50,000

GPU count

1

Quantization

FP16

Tokens/request

500 in, 500 out

Charts appear only when evidence existsUtilization defaults to 30%

Configure the workload

Change the inputs, then calculate a shareable result.

Hugging Face model

Utilization assumption

Defaulted to 30% so the page does not manufacture a self-hosting win.

Derived utilization

97%

Idle share paid

2.7%

Utilization assumption30%

A steady production service with real peaks and troughs.

At this level about 70% of every rented GPU hour sits idle, and it is billed at the same rate as a busy one.

Stay on API

The API is the lower-cost choice for this workload.

All 3 dated API comparisons supplied for this model reach the same recommendation at this workload. Provider-specific bills remain below.

Model

meta-llama/Llama-3.1-8B-Instruct

Hardware

1 x NVIDIA A10 24 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

16.1 GB
Model weights
1.1 GB
KV cache
19.7 GB
Total required
21.6 GB
Available
8,192 tokens
Served context used
2.6 GB
Runtime overhead
Exact transformer shape
KV cache method

Reproduce with vLLM

The exact command attached to this result.

vllm serve 'meta-llama/Llama-3.1-8B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9

Priced operating point

Utilization and idle share are shown before the token price.

Paid idle capacity

2.7%

158.9 of 5,844 billed GPU-hours sit idle. The provider still charges every hour.

Busy GPU-hours

5,685.1

Useful work and paid idle time sum to the billed total above.

97%
Derived utilization
2.7%
Idle share paid
37 output tok/s
Decode ceiling at concurrency 1
Physics ceiling

CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.

Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08

$5,844.00
Monthly GPU rent

$1.00/GPU-hour, OVHcloud, Public Cloud, on demand
Captured 2026-08-09 (today)

At configured-shape saturation, 500 input / 500 output tokens

$7.47 / 1M output tokens

At actual utilization

$7.68 / 1M output tokens

See where the money moves

Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.

Cost vs volume

Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.

Self-host, discrete pointsCheapest comparable API, discrete pointsCurrent traffic

Scroll horizontally to read every volume point.

500 requests per day, 7,609,375 output tokens per month: self-host $730.50, cheapest comparable API $0.46. 1,500 requests per day, 22,828,125 output tokens per month: self-host $730.50, cheapest comparable API $1.37. 5,000 requests per day, 76,093,750 output tokens per month: self-host $730.50, cheapest comparable API $4.57. 6,360 requests per day, 96,798,777 output tokens per month: self-host $730.50, cheapest comparable API $5.81. 6,489 requests per day, 98,754,308 output tokens per month: self-host $1,461.00, cheapest comparable API $5.93. 15,000 requests per day, 228,281,250 output tokens per month: self-host $2,191.50, cheapest comparable API $13.70. 50,000 requests per day, 760,937,500 output tokens per month: self-host $5,844.00, cheapest comparable API $45.66. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $17,532.00, cheapest comparable API $136.97. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $56,979.00, cheapest comparable API $456.56. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $170,937.00, cheapest comparable API $1,369.69. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $569,059.50, cheapest comparable API $4,565.63.

$569.1K$0500 req/day50,000 req/day5,000,000 req/day
Current self-host bill
$5,844.00
Current cheapest API
$45.66
Monthly output tokens
760,937,500
Source:
  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded, as of 2026-08-09
  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Groq, as of 2026-07-31
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting DeepInfra, as of 2026-07-31
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Novita AI, as of 2026-07-31
+What this chart means

Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.

Busy vs idle GPU-hours

5,844 billed GPU-hours, split by the engine into useful work and paid idle capacity.

Scroll horizontally to read every chart label and value.

Busy GPU-hours: 5,685.1, Engine-reported share doing token work. Idle GPU-hours billed: 158.9, Engine-reported share paid while idle.

  • Busy GPU-hours5,685.1
  • Idle GPU-hours billed158.9
Source:
  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded, as of 2026-08-09
  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08
+What this chart means

Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.

Provider price spread, NVIDIA A10 24 GB

USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.

Scroll horizontally to read every chart label and value.

Lambda Cloud: $1.29, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Microsoft Azure, East US 2: $3.20, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Oracle Cloud: $2.00, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest + Used: OVHcloud, Public Cloud: $1.00, on demand GPU-hour rate, captured 2026-08-09 (today), used by the engine for this calculation.

  • Lambda Cloud$1.29
  • Microsoft Azure, East US 2$3.20
  • Oracle Cloud$2.00
  • Cheapest + Used: OVHcloud, Public Cloud$1.00
Source:
  • Lambda Cloud, on demand, 1 day ago, captured 2026-08-08
  • Microsoft Azure, East US 2, on demand, 1 day ago, captured 2026-08-08
  • Oracle Cloud, on demand, 1 day ago, captured 2026-08-08
  • OVHcloud, Public Cloud, on demand, today, captured 2026-08-09
+What this chart means

Every rate is for the selected GPU and keeps its provider, tier, source, capture date, and age. A price without a date is not included.

Not measured on this config yet

No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.

Run this configuration

API comparisons

Input and output rates stay separate because tokenizers and workloads differ.

Groq

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 44,953,846,154 tokens

API monthly
$98.92
Self-host monthly
$5,844.00

Input: $0.05 / 1M tokens

Output: $0.08 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

DeepInfra

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 97,400,000,000 tokens

API monthly
$45.66
Self-host monthly
$5,844.00

Input: $0.02 / 1M tokens

Output: $0.04 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

Novita AI

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 83,485,714,286 tokens

API monthly
$53.27
Self-host monthly
$5,844.00

Input: $0.02 / 1M tokens

Output: $0.05 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA A10 datasheet

    As of 2026-08-08

  • Meta Llama 3.1 8B Hugging Face config

    As of 2026-08-08

  • RunInfra Engine feasibility runtime-overhead policy

    As of 2026-08-09

  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters

    As of 2026-08-08

  • vLLM serve CLI documentation

    As of 2026-08-08

  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded

    As of 2026-08-09

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Groq

    As of 2026-07-31

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting DeepInfra

    As of 2026-07-31

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Novita AI

    As of 2026-07-31

  • NVIDIA L4 product specification

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

Caveats

  • +Served context is 8192 tokens. It is a serving-capacity decision and is not inferred from average input or output tokens.
  • +KV cache uses the exact transformer-shape formula with 32 layers, 8 KV heads, and head dimension 128.
  • +Runtime overhead adds 2.57 GB to totalRequiredGb. This approximate 15% policy allowance covers CUDA context, activations, and vLLM overhead. Deployment-time profiling, including CUDA graph capture and allocator behavior, can differ.
  • +CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
  • +Compute prefill ceiling: 7783.059363799169 tokens/sec. Bandwidth decode ceiling: 37.35868494623601 tokens/sec.
  • +CalcInput has no GPU provider selector. The calculation uses the lowest supplied on-demand rate for A10: OVHcloud, Public Cloud at $1/GPU-hour.
  • +Replica count covers the configured average daily token work. Peak concurrency is absent from CalcInput, so additional replicas needed for burst traffic are not priced.
  • +The utilization control is a break-even capacity scenario. derivedUtilization is separately computed from the configured daily token work.
  • +API input and output prices are applied separately. Cross-provider token counts can still differ because tokenizers differ.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Stay on API

The API is the lower-cost choice for this workload.

All 3 dated API comparisons supplied for this model reach the same recommendation at this workload. Provider-specific bills remain below.

Model

meta-llama/Llama-3.1-8B-Instruct

Hardware

1 x NVIDIA A10 24 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

16.1 GB
Model weights
1.1 GB
KV cache
19.7 GB
Total required
21.6 GB
Available
8,192 tokens
Served context used
2.6 GB
Runtime overhead
Exact transformer shape
KV cache method

Reproduce with vLLM

The exact command attached to this result.

vllm serve 'meta-llama/Llama-3.1-8B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9

Priced operating point

Utilization and idle share are shown before the token price.

Paid idle capacity

2.7%

158.9 of 5,844 billed GPU-hours sit idle. The provider still charges every hour.

Busy GPU-hours

5,685.1

Useful work and paid idle time sum to the billed total above.

97%
Derived utilization
2.7%
Idle share paid
37 output tok/s
Decode ceiling at concurrency 1
Physics ceiling

CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.

Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08

$5,844.00
Monthly GPU rent

$1.00/GPU-hour, OVHcloud, Public Cloud, on demand
Captured 2026-08-09 (today)

At configured-shape saturation, 500 input / 500 output tokens

$7.47 / 1M output tokens

At actual utilization

$7.68 / 1M output tokens

See where the money moves

Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.

Cost vs volume

Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.

Self-host, discrete pointsCheapest comparable API, discrete pointsCurrent traffic

Scroll horizontally to read every volume point.

500 requests per day, 7,609,375 output tokens per month: self-host $730.50, cheapest comparable API $0.46. 1,500 requests per day, 22,828,125 output tokens per month: self-host $730.50, cheapest comparable API $1.37. 5,000 requests per day, 76,093,750 output tokens per month: self-host $730.50, cheapest comparable API $4.57. 6,360 requests per day, 96,798,777 output tokens per month: self-host $730.50, cheapest comparable API $5.81. 6,489 requests per day, 98,754,308 output tokens per month: self-host $1,461.00, cheapest comparable API $5.93. 15,000 requests per day, 228,281,250 output tokens per month: self-host $2,191.50, cheapest comparable API $13.70. 50,000 requests per day, 760,937,500 output tokens per month: self-host $5,844.00, cheapest comparable API $45.66. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $17,532.00, cheapest comparable API $136.97. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $56,979.00, cheapest comparable API $456.56. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $170,937.00, cheapest comparable API $1,369.69. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $569,059.50, cheapest comparable API $4,565.63.

$569.1K$0500 req/day50,000 req/day5,000,000 req/day
Current self-host bill
$5,844.00
Current cheapest API
$45.66
Monthly output tokens
760,937,500
Source:
  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded, as of 2026-08-09
  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Groq, as of 2026-07-31
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting DeepInfra, as of 2026-07-31
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Novita AI, as of 2026-07-31
+What this chart means

Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.

Busy vs idle GPU-hours

5,844 billed GPU-hours, split by the engine into useful work and paid idle capacity.

Scroll horizontally to read every chart label and value.

Busy GPU-hours: 5,685.1, Engine-reported share doing token work. Idle GPU-hours billed: 158.9, Engine-reported share paid while idle.

  • Busy GPU-hours5,685.1
  • Idle GPU-hours billed158.9
Source:
  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded, as of 2026-08-09
  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08
+What this chart means

Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.

Provider price spread, NVIDIA A10 24 GB

USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.

Scroll horizontally to read every chart label and value.

Lambda Cloud: $1.29, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Microsoft Azure, East US 2: $3.20, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Oracle Cloud: $2.00, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest + Used: OVHcloud, Public Cloud: $1.00, on demand GPU-hour rate, captured 2026-08-09 (today), used by the engine for this calculation.

  • Lambda Cloud$1.29
  • Microsoft Azure, East US 2$3.20
  • Oracle Cloud$2.00
  • Cheapest + Used: OVHcloud, Public Cloud$1.00
Source:
  • Lambda Cloud, on demand, 1 day ago, captured 2026-08-08
  • Microsoft Azure, East US 2, on demand, 1 day ago, captured 2026-08-08
  • Oracle Cloud, on demand, 1 day ago, captured 2026-08-08
  • OVHcloud, Public Cloud, on demand, today, captured 2026-08-09
+What this chart means

Every rate is for the selected GPU and keeps its provider, tier, source, capture date, and age. A price without a date is not included.

Not measured on this config yet

No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.

Run this configuration

API comparisons

Input and output rates stay separate because tokenizers and workloads differ.

Groq

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 44,953,846,154 tokens

API monthly
$98.92
Self-host monthly
$5,844.00

Input: $0.05 / 1M tokens

Output: $0.08 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

DeepInfra

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 97,400,000,000 tokens

API monthly
$45.66
Self-host monthly
$5,844.00

Input: $0.02 / 1M tokens

Output: $0.04 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

Novita AI

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 83,485,714,286 tokens

API monthly
$53.27
Self-host monthly
$5,844.00

Input: $0.02 / 1M tokens

Output: $0.05 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA A10 datasheet

    As of 2026-08-08

  • Meta Llama 3.1 8B Hugging Face config

    As of 2026-08-08

  • RunInfra Engine feasibility runtime-overhead policy

    As of 2026-08-09

  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters

    As of 2026-08-08

  • vLLM serve CLI documentation

    As of 2026-08-08

  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded

    As of 2026-08-09

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Groq

    As of 2026-07-31

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting DeepInfra

    As of 2026-07-31

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Novita AI

    As of 2026-07-31

  • NVIDIA L4 product specification

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

Caveats

  • +Served context is 8192 tokens. It is a serving-capacity decision and is not inferred from average input or output tokens.
  • +KV cache uses the exact transformer-shape formula with 32 layers, 8 KV heads, and head dimension 128.
  • +Runtime overhead adds 2.57 GB to totalRequiredGb. This approximate 15% policy allowance covers CUDA context, activations, and vLLM overhead. Deployment-time profiling, including CUDA graph capture and allocator behavior, can differ.
  • +CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
  • +Compute prefill ceiling: 7783.059363799169 tokens/sec. Bandwidth decode ceiling: 37.35868494623601 tokens/sec.
  • +CalcInput has no GPU provider selector. The calculation uses the lowest supplied on-demand rate for A10: OVHcloud, Public Cloud at $1/GPU-hour.
  • +Replica count covers the configured average daily token work. Peak concurrency is absent from CalcInput, so additional replicas needed for burst traffic are not priced.
  • +The utilization control is a break-even capacity scenario. derivedUtilization is separately computed from the configured daily token work.
  • +API input and output prices are applied separately. Cross-provider token counts can still differ because tokenizers differ.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Stay on API

The API is the lower-cost choice for this workload.

All 3 dated API comparisons supplied for this model reach the same recommendation at this workload. Provider-specific bills remain below.

Model

meta-llama/Llama-3.1-8B-Instruct

Hardware

1 x NVIDIA A10 24 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

16.1 GB
Model weights
1.1 GB
KV cache
19.7 GB
Total required
21.6 GB
Available
8,192 tokens
Served context used
2.6 GB
Runtime overhead
Exact transformer shape
KV cache method

Reproduce with vLLM

The exact command attached to this result.

vllm serve 'meta-llama/Llama-3.1-8B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9

Priced operating point

Utilization and idle share are shown before the token price.

Paid idle capacity

2.7%

158.9 of 5,844 billed GPU-hours sit idle. The provider still charges every hour.

Busy GPU-hours

5,685.1

Useful work and paid idle time sum to the billed total above.

97%
Derived utilization
2.7%
Idle share paid
37 output tok/s
Decode ceiling at concurrency 1
Physics ceiling

CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.

Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08

$5,844.00
Monthly GPU rent

$1.00/GPU-hour, OVHcloud, Public Cloud, on demand
Captured 2026-08-09 (today)

At configured-shape saturation, 500 input / 500 output tokens

$7.47 / 1M output tokens

At actual utilization

$7.68 / 1M output tokens

See where the money moves

Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.

Cost vs volume

Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.

Self-host, discrete pointsCheapest comparable API, discrete pointsCurrent traffic

Scroll horizontally to read every volume point.

500 requests per day, 7,609,375 output tokens per month: self-host $730.50, cheapest comparable API $0.46. 1,500 requests per day, 22,828,125 output tokens per month: self-host $730.50, cheapest comparable API $1.37. 5,000 requests per day, 76,093,750 output tokens per month: self-host $730.50, cheapest comparable API $4.57. 6,360 requests per day, 96,798,777 output tokens per month: self-host $730.50, cheapest comparable API $5.81. 6,489 requests per day, 98,754,308 output tokens per month: self-host $1,461.00, cheapest comparable API $5.93. 15,000 requests per day, 228,281,250 output tokens per month: self-host $2,191.50, cheapest comparable API $13.70. 50,000 requests per day, 760,937,500 output tokens per month: self-host $5,844.00, cheapest comparable API $45.66. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $17,532.00, cheapest comparable API $136.97. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $56,979.00, cheapest comparable API $456.56. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $170,937.00, cheapest comparable API $1,369.69. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $569,059.50, cheapest comparable API $4,565.63.

$569.1K$0500 req/day50,000 req/day5,000,000 req/day
Current self-host bill
$5,844.00
Current cheapest API
$45.66
Monthly output tokens
760,937,500
Source:
  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded, as of 2026-08-09
  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Groq, as of 2026-07-31
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting DeepInfra, as of 2026-07-31
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Novita AI, as of 2026-07-31
+What this chart means

Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.

Busy vs idle GPU-hours

5,844 billed GPU-hours, split by the engine into useful work and paid idle capacity.

Scroll horizontally to read every chart label and value.

Busy GPU-hours: 5,685.1, Engine-reported share doing token work. Idle GPU-hours billed: 158.9, Engine-reported share paid while idle.

  • Busy GPU-hours5,685.1
  • Idle GPU-hours billed158.9
Source:
  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded, as of 2026-08-09
  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08
+What this chart means

Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.

Provider price spread, NVIDIA A10 24 GB

USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.

Scroll horizontally to read every chart label and value.

Lambda Cloud: $1.29, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Microsoft Azure, East US 2: $3.20, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Oracle Cloud: $2.00, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest + Used: OVHcloud, Public Cloud: $1.00, on demand GPU-hour rate, captured 2026-08-09 (today), used by the engine for this calculation.

  • Lambda Cloud$1.29
  • Microsoft Azure, East US 2$3.20
  • Oracle Cloud$2.00
  • Cheapest + Used: OVHcloud, Public Cloud$1.00
Source:
  • Lambda Cloud, on demand, 1 day ago, captured 2026-08-08
  • Microsoft Azure, East US 2, on demand, 1 day ago, captured 2026-08-08
  • Oracle Cloud, on demand, 1 day ago, captured 2026-08-08
  • OVHcloud, Public Cloud, on demand, today, captured 2026-08-09
+What this chart means

Every rate is for the selected GPU and keeps its provider, tier, source, capture date, and age. A price without a date is not included.

Not measured on this config yet

No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.

Run this configuration

API comparisons

Input and output rates stay separate because tokenizers and workloads differ.

Groq

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 44,953,846,154 tokens

API monthly
$98.92
Self-host monthly
$5,844.00

Input: $0.05 / 1M tokens

Output: $0.08 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

DeepInfra

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 97,400,000,000 tokens

API monthly
$45.66
Self-host monthly
$5,844.00

Input: $0.02 / 1M tokens

Output: $0.04 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

Novita AI

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 83,485,714,286 tokens

API monthly
$53.27
Self-host monthly
$5,844.00

Input: $0.02 / 1M tokens

Output: $0.05 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA A10 datasheet

    As of 2026-08-08

  • Meta Llama 3.1 8B Hugging Face config

    As of 2026-08-08

  • RunInfra Engine feasibility runtime-overhead policy

    As of 2026-08-09

  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters

    As of 2026-08-08

  • vLLM serve CLI documentation

    As of 2026-08-08

  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded

    As of 2026-08-09

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Groq

    As of 2026-07-31

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting DeepInfra

    As of 2026-07-31

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Novita AI

    As of 2026-07-31

  • NVIDIA L4 product specification

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

Caveats

  • +Served context is 8192 tokens. It is a serving-capacity decision and is not inferred from average input or output tokens.
  • +KV cache uses the exact transformer-shape formula with 32 layers, 8 KV heads, and head dimension 128.
  • +Runtime overhead adds 2.57 GB to totalRequiredGb. This approximate 15% policy allowance covers CUDA context, activations, and vLLM overhead. Deployment-time profiling, including CUDA graph capture and allocator behavior, can differ.
  • +CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
  • +Compute prefill ceiling: 7783.059363799169 tokens/sec. Bandwidth decode ceiling: 37.35868494623601 tokens/sec.
  • +CalcInput has no GPU provider selector. The calculation uses the lowest supplied on-demand rate for A10: OVHcloud, Public Cloud at $1/GPU-hour.
  • +Replica count covers the configured average daily token work. Peak concurrency is absent from CalcInput, so additional replicas needed for burst traffic are not priced.
  • +The utilization control is a break-even capacity scenario. derivedUtilization is separately computed from the configured daily token work.
  • +API input and output prices are applied separately. Cross-provider token counts can still differ because tokenizers differ.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Stay on API

The API is the lower-cost choice for this workload.

All 3 dated API comparisons supplied for this model reach the same recommendation at this workload. Provider-specific bills remain below.

Model

meta-llama/Llama-3.1-8B-Instruct

Hardware

1 x NVIDIA A10 24 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

16.1 GB
Model weights
1.1 GB
KV cache
19.7 GB
Total required
21.6 GB
Available
8,192 tokens
Served context used
2.6 GB
Runtime overhead
Exact transformer shape
KV cache method

Reproduce with vLLM

The exact command attached to this result.

vllm serve 'meta-llama/Llama-3.1-8B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9

Priced operating point

Utilization and idle share are shown before the token price.

Paid idle capacity

2.7%

158.9 of 5,844 billed GPU-hours sit idle. The provider still charges every hour.

Busy GPU-hours

5,685.1

Useful work and paid idle time sum to the billed total above.

97%
Derived utilization
2.7%
Idle share paid
37 output tok/s
Decode ceiling at concurrency 1
Physics ceiling

CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.

Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, captured 2026-08-08

$5,844.00
Monthly GPU rent

$1.00/GPU-hour, OVHcloud, Public Cloud, on demand
Captured 2026-08-09 (today)

At configured-shape saturation, 500 input / 500 output tokens

$7.47 / 1M output tokens

At actual utilization

$7.68 / 1M output tokens

See where the money moves

Each available view stays on its own evidence-backed scale. Missing rows remain visibly absent.

Cost vs volume

Each marker is a full engine recomputation. Self-host prices are never connected or interpolated.

Self-host, discrete pointsCheapest comparable API, discrete pointsCurrent traffic

Scroll horizontally to read every volume point.

500 requests per day, 7,609,375 output tokens per month: self-host $730.50, cheapest comparable API $0.46. 1,500 requests per day, 22,828,125 output tokens per month: self-host $730.50, cheapest comparable API $1.37. 5,000 requests per day, 76,093,750 output tokens per month: self-host $730.50, cheapest comparable API $4.57. 6,360 requests per day, 96,798,777 output tokens per month: self-host $730.50, cheapest comparable API $5.81. 6,489 requests per day, 98,754,308 output tokens per month: self-host $1,461.00, cheapest comparable API $5.93. 15,000 requests per day, 228,281,250 output tokens per month: self-host $2,191.50, cheapest comparable API $13.70. 50,000 requests per day, 760,937,500 output tokens per month: self-host $5,844.00, cheapest comparable API $45.66. 150,000 requests per day, 2,282,812,500 output tokens per month: self-host $17,532.00, cheapest comparable API $136.97. 500,000 requests per day, 7,609,375,000 output tokens per month: self-host $56,979.00, cheapest comparable API $456.56. 1,500,000 requests per day, 22,828,125,000 output tokens per month: self-host $170,937.00, cheapest comparable API $1,369.69. 5,000,000 requests per day, 76,093,750,000 output tokens per month: self-host $569,059.50, cheapest comparable API $4,565.63.

$569.1K$0500 req/day50,000 req/day5,000,000 req/day
Current self-host bill
$5,844.00
Current cheapest API
$45.66
Monthly output tokens
760,937,500
Source:
  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded, as of 2026-08-09
  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Groq, as of 2026-07-31
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting DeepInfra, as of 2026-07-31
  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Novita AI, as of 2026-07-31
+What this chart means

Replica counts change in whole units. Isolated self-host markers preserve each emitted price without drawing a smooth, straight, or step transition at a volume the engine did not emit. Missing API values produce no marker.

Busy vs idle GPU-hours

5,844 billed GPU-hours, split by the engine into useful work and paid idle capacity.

Scroll horizontally to read every chart label and value.

Busy GPU-hours: 5,685.1, Engine-reported share doing token work. Idle GPU-hours billed: 158.9, Engine-reported share paid while idle.

  • Busy GPU-hours5,685.1
  • Idle GPU-hours billed158.9
Source:
  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded, as of 2026-08-09
  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08
+What this chart means

Busy and idle hours are emitted by the engine and sum to the billed GPU-hours. The provider charges both, so the idle share remains visible beside the monthly rent.

Provider price spread, NVIDIA A10 24 GB

USD per GPU-hour. Cheapest captured offers and the rate used in this calculation are marked separately.

Scroll horizontally to read every chart label and value.

Lambda Cloud: $1.29, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Microsoft Azure, East US 2: $3.20, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Oracle Cloud: $2.00, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest + Used: OVHcloud, Public Cloud: $1.00, on demand GPU-hour rate, captured 2026-08-09 (today), used by the engine for this calculation.

  • Lambda Cloud$1.29
  • Microsoft Azure, East US 2$3.20
  • Oracle Cloud$2.00
  • Cheapest + Used: OVHcloud, Public Cloud$1.00
Source:
  • Lambda Cloud, on demand, 1 day ago, captured 2026-08-08
  • Microsoft Azure, East US 2, on demand, 1 day ago, captured 2026-08-08
  • Oracle Cloud, on demand, 1 day ago, captured 2026-08-08
  • OVHcloud, Public Cloud, on demand, today, captured 2026-08-09
+What this chart means

Every rate is for the selected GPU and keeps its provider, tier, source, capture date, and age. A price without a date is not included.

Not measured on this config yet

No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.

Run this configuration

API comparisons

Input and output rates stay separate because tokenizers and workloads differ.

Groq

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 44,953,846,154 tokens

API monthly
$98.92
Self-host monthly
$5,844.00

Input: $0.05 / 1M tokens

Output: $0.08 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

DeepInfra

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 97,400,000,000 tokens

API monthly
$45.66
Self-host monthly
$5,844.00

Input: $0.02 / 1M tokens

Output: $0.04 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

Novita AI

The API is the lower-cost choice for this workload.

That recommendation compares the two monthly bills at the configured volume. Physical capacity to reach break-even is reported separately.

Capacity ceiling

Can this hardware reach API break-even at the selected utilization? No

Required monthly output volume: 83,485,714,286 tokens

API monthly
$53.27
Self-host monthly
$5,844.00

Input: $0.02 / 1M tokens

Output: $0.05 / 1M tokens

Captured 2026-07-31 (9 days ago)

Tokenizers differ across providers, so per-token prices are not a like-for-like ranking of identical text.

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA A10 datasheet

    As of 2026-08-08

  • Meta Llama 3.1 8B Hugging Face config

    As of 2026-08-08

  • RunInfra Engine feasibility runtime-overhead policy

    As of 2026-08-09

  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters

    As of 2026-08-08

  • vLLM serve CLI documentation

    As of 2026-08-08

  • OVHcloud, Public Cloud pricing, a10-45 one-GPU A10 24 GB hourly consumption, France Gravelines, tax excluded

    As of 2026-08-09

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Groq

    As of 2026-07-31

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting DeepInfra

    As of 2026-07-31

  • RightNow AI inference-cost-truth (CC BY 4.0), quoting Novita AI

    As of 2026-07-31

  • NVIDIA L4 product specification

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

Caveats

  • +Served context is 8192 tokens. It is a serving-capacity decision and is not inferred from average input or output tokens.
  • +KV cache uses the exact transformer-shape formula with 32 layers, 8 KV heads, and head dimension 128.
  • +Runtime overhead adds 2.57 GB to totalRequiredGb. This approximate 15% policy allowance covers CUDA context, activations, and vLLM overhead. Deployment-time profiling, including CUDA graph capture and allocator behavior, can differ.
  • +CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
  • +Compute prefill ceiling: 7783.059363799169 tokens/sec. Bandwidth decode ceiling: 37.35868494623601 tokens/sec.
  • +CalcInput has no GPU provider selector. The calculation uses the lowest supplied on-demand rate for A10: OVHcloud, Public Cloud at $1/GPU-hour.
  • +Replica count covers the configured average daily token work. Peak concurrency is absent from CalcInput, so additional replicas needed for burst traffic are not priced.
  • +The utilization control is a break-even capacity scenario. derivedUtilization is separately computed from the configured daily token work.
  • +API input and output prices are applied separately. Cross-provider token counts can still differ because tokenizers differ.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Share this calculation

The badge keeps its evidence label and links to the full result, sources, and caveats.

Calculated by RunInfra

HTML

<a href="https://runinfra.ai/calc/llama-3.1-8b/a10/50k-req-day?context=8192"><img src="https://runinfra.ai/api/calc/badge?hfId=meta-llama%2FLlama-3.1-8B-Instruct&amp;gpuId=A10&amp;gpuCount=1&amp;requestsPerDay=50000&amp;avgInputTokens=500&amp;avgOutputTokens=500&amp;utilization=0.3&amp;quantization=fp16" alt="Calculated by RunInfra"></a>

Markdown

[![Calculated by RunInfra](https://runinfra.ai/api/calc/badge?hfId=meta-llama%2FLlama-3.1-8B-Instruct&gpuId=A10&gpuCount=1&requestsPerDay=50000&avgInputTokens=500&avgOutputTokens=500&utilization=0.3&quantization=fp16)](https://runinfra.ai/calc/llama-3.1-8b/a10/50k-req-day?context=8192)

Use it from an agent or script

GET the current inputs, or inspect the descriptor for accepted parameters and response fields.

Endpoint
GET https://runinfra.ai/api/calchttps://runinfra.ai/api/calc?hfId=meta-llama%2FLlama-3.1-8B-Instruct&gpuId=A10&gpuCount=1&requestsPerDay=50000&avgInputTokens=500&avgOutputTokens=500&utilization=0.3&quantization=fp16
Canonical result
https://runinfra.ai/calc/llama-3.1-8b/a10/50k-req-day?context=8192
Descriptor
https://runinfra.ai/api/calc/schema

28 GPU ids, 7 quantizations

Turn this exact configuration into a deployment, describe what you need

Describe how you want to deploy this configuration...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy