RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

See what this model actually costs to serve.

Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.

gemma-2-9b-it does not fit on 1 x NVIDIA A10 24 GB: 26.4 GB required, 21.6 GB available.

Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.

Model

google/gemma-2-9b-it

GPU

NVIDIA A10 24 GB

Requests/day

50,000

GPU count

1

Quantization

FP16

Tokens/request

500 in, 500 out

Charts appear only when evidence existsUtilization defaults to 30%

Configure the workload

Change the inputs, then calculate a shareable result.

Hugging Face model

Utilization assumption

Defaulted to 30% so the page does not manufacture a self-hosting win.

Derived utilization

Needs throughput

Idle share paid

Needs throughput

Utilization assumption30%

A steady production service with real peaks and troughs.

At this level about 70% of every rented GPU hour sits idle, and it is billed at the same rate as a busy one.

Does not fit

This configuration needs more memory.

The model and KV cache need more VRAM than the selected GPU count provides. Cost is intentionally hidden for a configuration that cannot run.

Model

google/gemma-2-9b-it

Hardware

1 x NVIDIA A10 24 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

18.5 GB
Model weights
4.4 GB
KV cache
26.4 GB
Total required
21.6 GB
Available
8,192 tokens
Served context used
3.4 GB
Runtime overhead
Parameter-only heuristic
KV cache method

Memory-fit rental floors

Published rent floors for memory-fitting configurations.

Each row uses an engine GPU fit when present; the remaining catalog GPUs apply this result's required memory to published VRAM. GPU count rises only enough to clear memory.

Engine memory-fit result

Memory fit: 1x NVIDIA GeForce RTX 5090 32 GB at 90% vLLM GPU memory utilization.

On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The rentable GPU catalog is ranked by one-deployment monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. These are arithmetic floors, not provider offers. Throughput, replica demand, traffic capacity, topology support, and current availability are not established.

  1. Floor rank 1

    2 x NVIDIA RTX A4000 16 GB

    Memory fits

    26.4 GB required, 28.8 GB available inside the engine memory budget.

    Current result memory total, Ampere, 448 GB/s memory bandwidth

    One-deployment monthly rent floor

    $219.15

    Hyperstack, on demand

    Captured 2026-08-08

    Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source
  2. Floor rank 2

    2 x NVIDIA RTX A5000 24 GB

    Memory fits

    26.4 GB required, 43.2 GB available inside the engine memory budget.

    Current result memory total, Ampere, 768 GB/s memory bandwidth

    One-deployment monthly rent floor

    $233.76

    RunPod, Community Cloud, on demand

    Captured 2026-08-08

    Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source
  3. Floor rank 3

    1 x NVIDIA RTX A6000 48 GB

    Memory fits

    26.4 GB required, 43.2 GB available inside the engine memory budget.

    Current result memory total, Ampere, 768 GB/s memory bandwidth

    One-deployment monthly rent floor

    $241.07

    RunPod, Community Cloud, on demand

    Captured 2026-08-08

    Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source

What to change

Computed actions and cost signals appear only when the result supports them.

RecommendationComputed memory fit

Consider FP8 or INT8 if the checkpoint and hardware support it

FP16 weights use 18.5 GB and bring total memory to 26.4 GB; FP8 or INT8 would lower weights to 9.2 GB and total memory to 13.7 GB, which fits inside 21.6 GB.

Show 2 more computed options
RecommendationComputed memory fit

Reduce context length or serving concurrency

Weights use 18.5 GB and fit inside 21.6 GB, but the 4.4 GB KV cache raises total memory to 26.4 GB; reduce --max-model-len or peak concurrency before adding hardware.

RecommendationComputed memory fit

Raise tensor parallel size to 2

The current vLLM command uses tensor parallel size 1 and provides 21.6 GB, while 26.4 GB requires 2 same-GPU workers for 43.2 GB; raise --tensor-parallel-size to 2 if the model topology supports it.

Decode bandwidth ceiling

A physics bound, not an expected throughput result.

32 output tok/sCeiling

CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.

Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08

Reproduce with vLLM

The exact command attached to this result.

vllm serve 'google/gemma-2-9b-it' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA A10 datasheet

    As of 2026-08-08

  • RunInfra Engine parameter-only KV-cache heuristic

    As of 2026-08-08

  • RunInfra Engine feasibility runtime-overhead policy

    As of 2026-08-09

  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters

    As of 2026-08-08

  • vLLM serve CLI documentation

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

  • NVIDIA H100 Tensor Core GPU datasheet

    As of 2026-08-08

Caveats

  • +Served context is 8192 tokens. It is a serving-capacity decision and is not inferred from average input or output tokens.
  • +Resolved layer count, KV-head count, or head dimension is absent. KV cache uses the parameter-only RunInfra Engine heuristic rather than an exact architecture calculation.
  • +Runtime overhead adds 3.44 GB to totalRequiredGb. This approximate 15% policy allowance covers CUDA context, activations, and vLLM overhead. Deployment-time profiling, including CUDA graph capture and allocator behavior, can differ.
  • +CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
  • +Compute prefill ceiling: 6762.82064244471 tokens/sec. Bandwidth decode ceiling: 32.46153908373461 tokens/sec.
  • +CalcInput has no GPU provider selector. The calculation uses the lowest supplied on-demand rate for A10: OVHcloud, Public Cloud at $1/GPU-hour.
  • +The selected configuration does not fit, so no cost is returned.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Does not fit

This configuration needs more memory.

The model and KV cache need more VRAM than the selected GPU count provides. Cost is intentionally hidden for a configuration that cannot run.

Model

google/gemma-2-9b-it

Hardware

1 x NVIDIA A10 24 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

18.5 GB
Model weights
4.4 GB
KV cache
26.4 GB
Total required
21.6 GB
Available
8,192 tokens
Served context used
3.4 GB
Runtime overhead
Parameter-only heuristic
KV cache method

Memory-fit rental floors

Published rent floors for memory-fitting configurations.

Each row uses an engine GPU fit when present; the remaining catalog GPUs apply this result's required memory to published VRAM. GPU count rises only enough to clear memory.

Engine memory-fit result

Memory fit: 1x NVIDIA GeForce RTX 5090 32 GB at 90% vLLM GPU memory utilization.

On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The rentable GPU catalog is ranked by one-deployment monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. These are arithmetic floors, not provider offers. Throughput, replica demand, traffic capacity, topology support, and current availability are not established.

  1. Floor rank 1

    2 x NVIDIA RTX A4000 16 GB

    Memory fits

    26.4 GB required, 28.8 GB available inside the engine memory budget.

    Current result memory total, Ampere, 448 GB/s memory bandwidth

    One-deployment monthly rent floor

    $219.15

    Hyperstack, on demand

    Captured 2026-08-08

    Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source
  2. Floor rank 2

    2 x NVIDIA RTX A5000 24 GB

    Memory fits

    26.4 GB required, 43.2 GB available inside the engine memory budget.

    Current result memory total, Ampere, 768 GB/s memory bandwidth

    One-deployment monthly rent floor

    $233.76

    RunPod, Community Cloud, on demand

    Captured 2026-08-08

    Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source
  3. Floor rank 3

    1 x NVIDIA RTX A6000 48 GB

    Memory fits

    26.4 GB required, 43.2 GB available inside the engine memory budget.

    Current result memory total, Ampere, 768 GB/s memory bandwidth

    One-deployment monthly rent floor

    $241.07

    RunPod, Community Cloud, on demand

    Captured 2026-08-08

    Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source

What to change

Computed actions and cost signals appear only when the result supports them.

RecommendationComputed memory fit

Consider FP8 or INT8 if the checkpoint and hardware support it

FP16 weights use 18.5 GB and bring total memory to 26.4 GB; FP8 or INT8 would lower weights to 9.2 GB and total memory to 13.7 GB, which fits inside 21.6 GB.

Show 2 more computed options
RecommendationComputed memory fit

Reduce context length or serving concurrency

Weights use 18.5 GB and fit inside 21.6 GB, but the 4.4 GB KV cache raises total memory to 26.4 GB; reduce --max-model-len or peak concurrency before adding hardware.

RecommendationComputed memory fit

Raise tensor parallel size to 2

The current vLLM command uses tensor parallel size 1 and provides 21.6 GB, while 26.4 GB requires 2 same-GPU workers for 43.2 GB; raise --tensor-parallel-size to 2 if the model topology supports it.

Decode bandwidth ceiling

A physics bound, not an expected throughput result.

32 output tok/sCeiling

CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.

Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08

Reproduce with vLLM

The exact command attached to this result.

vllm serve 'google/gemma-2-9b-it' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA A10 datasheet

    As of 2026-08-08

  • RunInfra Engine parameter-only KV-cache heuristic

    As of 2026-08-08

  • RunInfra Engine feasibility runtime-overhead policy

    As of 2026-08-09

  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters

    As of 2026-08-08

  • vLLM serve CLI documentation

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

  • NVIDIA H100 Tensor Core GPU datasheet

    As of 2026-08-08

Caveats

  • +Served context is 8192 tokens. It is a serving-capacity decision and is not inferred from average input or output tokens.
  • +Resolved layer count, KV-head count, or head dimension is absent. KV cache uses the parameter-only RunInfra Engine heuristic rather than an exact architecture calculation.
  • +Runtime overhead adds 3.44 GB to totalRequiredGb. This approximate 15% policy allowance covers CUDA context, activations, and vLLM overhead. Deployment-time profiling, including CUDA graph capture and allocator behavior, can differ.
  • +CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
  • +Compute prefill ceiling: 6762.82064244471 tokens/sec. Bandwidth decode ceiling: 32.46153908373461 tokens/sec.
  • +CalcInput has no GPU provider selector. The calculation uses the lowest supplied on-demand rate for A10: OVHcloud, Public Cloud at $1/GPU-hour.
  • +The selected configuration does not fit, so no cost is returned.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Does not fit

This configuration needs more memory.

The model and KV cache need more VRAM than the selected GPU count provides. Cost is intentionally hidden for a configuration that cannot run.

Model

google/gemma-2-9b-it

Hardware

1 x NVIDIA A10 24 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

18.5 GB
Model weights
4.4 GB
KV cache
26.4 GB
Total required
21.6 GB
Available
8,192 tokens
Served context used
3.4 GB
Runtime overhead
Parameter-only heuristic
KV cache method

Memory-fit rental floors

Published rent floors for memory-fitting configurations.

Each row uses an engine GPU fit when present; the remaining catalog GPUs apply this result's required memory to published VRAM. GPU count rises only enough to clear memory.

Engine memory-fit result

Memory fit: 1x NVIDIA GeForce RTX 5090 32 GB at 90% vLLM GPU memory utilization.

On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The rentable GPU catalog is ranked by one-deployment monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. These are arithmetic floors, not provider offers. Throughput, replica demand, traffic capacity, topology support, and current availability are not established.

  1. Floor rank 1

    2 x NVIDIA RTX A4000 16 GB

    Memory fits

    26.4 GB required, 28.8 GB available inside the engine memory budget.

    Current result memory total, Ampere, 448 GB/s memory bandwidth

    One-deployment monthly rent floor

    $219.15

    Hyperstack, on demand

    Captured 2026-08-08

    Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source
  2. Floor rank 2

    2 x NVIDIA RTX A5000 24 GB

    Memory fits

    26.4 GB required, 43.2 GB available inside the engine memory budget.

    Current result memory total, Ampere, 768 GB/s memory bandwidth

    One-deployment monthly rent floor

    $233.76

    RunPod, Community Cloud, on demand

    Captured 2026-08-08

    Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source
  3. Floor rank 3

    1 x NVIDIA RTX A6000 48 GB

    Memory fits

    26.4 GB required, 43.2 GB available inside the engine memory budget.

    Current result memory total, Ampere, 768 GB/s memory bandwidth

    One-deployment monthly rent floor

    $241.07

    RunPod, Community Cloud, on demand

    Captured 2026-08-08

    Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source

What to change

Computed actions and cost signals appear only when the result supports them.

RecommendationComputed memory fit

Consider FP8 or INT8 if the checkpoint and hardware support it

FP16 weights use 18.5 GB and bring total memory to 26.4 GB; FP8 or INT8 would lower weights to 9.2 GB and total memory to 13.7 GB, which fits inside 21.6 GB.

Show 2 more computed options
RecommendationComputed memory fit

Reduce context length or serving concurrency

Weights use 18.5 GB and fit inside 21.6 GB, but the 4.4 GB KV cache raises total memory to 26.4 GB; reduce --max-model-len or peak concurrency before adding hardware.

RecommendationComputed memory fit

Raise tensor parallel size to 2

The current vLLM command uses tensor parallel size 1 and provides 21.6 GB, while 26.4 GB requires 2 same-GPU workers for 43.2 GB; raise --tensor-parallel-size to 2 if the model topology supports it.

Decode bandwidth ceiling

A physics bound, not an expected throughput result.

32 output tok/sCeiling

CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.

Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08

Reproduce with vLLM

The exact command attached to this result.

vllm serve 'google/gemma-2-9b-it' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA A10 datasheet

    As of 2026-08-08

  • RunInfra Engine parameter-only KV-cache heuristic

    As of 2026-08-08

  • RunInfra Engine feasibility runtime-overhead policy

    As of 2026-08-09

  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters

    As of 2026-08-08

  • vLLM serve CLI documentation

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

  • NVIDIA H100 Tensor Core GPU datasheet

    As of 2026-08-08

Caveats

  • +Served context is 8192 tokens. It is a serving-capacity decision and is not inferred from average input or output tokens.
  • +Resolved layer count, KV-head count, or head dimension is absent. KV cache uses the parameter-only RunInfra Engine heuristic rather than an exact architecture calculation.
  • +Runtime overhead adds 3.44 GB to totalRequiredGb. This approximate 15% policy allowance covers CUDA context, activations, and vLLM overhead. Deployment-time profiling, including CUDA graph capture and allocator behavior, can differ.
  • +CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
  • +Compute prefill ceiling: 6762.82064244471 tokens/sec. Bandwidth decode ceiling: 32.46153908373461 tokens/sec.
  • +CalcInput has no GPU provider selector. The calculation uses the lowest supplied on-demand rate for A10: OVHcloud, Public Cloud at $1/GPU-hour.
  • +The selected configuration does not fit, so no cost is returned.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Does not fit

This configuration needs more memory.

The model and KV cache need more VRAM than the selected GPU count provides. Cost is intentionally hidden for a configuration that cannot run.

Model

google/gemma-2-9b-it

Hardware

1 x NVIDIA A10 24 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

18.5 GB
Model weights
4.4 GB
KV cache
26.4 GB
Total required
21.6 GB
Available
8,192 tokens
Served context used
3.4 GB
Runtime overhead
Parameter-only heuristic
KV cache method

Memory-fit rental floors

Published rent floors for memory-fitting configurations.

Each row uses an engine GPU fit when present; the remaining catalog GPUs apply this result's required memory to published VRAM. GPU count rises only enough to clear memory.

Engine memory-fit result

Memory fit: 1x NVIDIA GeForce RTX 5090 32 GB at 90% vLLM GPU memory utilization.

On-demand rates are preferred for each GPU; when none is captured, the lowest dated captured tier is used. The rentable GPU catalog is ranked by one-deployment monthly rent over the contract's 730.5-hour month, rounded to cents. Equal displayed prices share a rank. These are arithmetic floors, not provider offers. Throughput, replica demand, traffic capacity, topology support, and current availability are not established.

  1. Floor rank 1

    2 x NVIDIA RTX A4000 16 GB

    Memory fits

    26.4 GB required, 28.8 GB available inside the engine memory budget.

    Current result memory total, Ampere, 448 GB/s memory bandwidth

    One-deployment monthly rent floor

    $219.15

    Hyperstack, on demand

    Captured 2026-08-08

    Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source
  2. Floor rank 2

    2 x NVIDIA RTX A5000 24 GB

    Memory fits

    26.4 GB required, 43.2 GB available inside the engine memory budget.

    Current result memory total, Ampere, 768 GB/s memory bandwidth

    One-deployment monthly rent floor

    $233.76

    RunPod, Community Cloud, on demand

    Captured 2026-08-08

    Arithmetic floor for 2 GPUs over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source
  3. Floor rank 3

    1 x NVIDIA RTX A6000 48 GB

    Memory fits

    26.4 GB required, 43.2 GB available inside the engine memory budget.

    Current result memory total, Ampere, 768 GB/s memory bandwidth

    One-deployment monthly rent floor

    $241.07

    RunPod, Community Cloud, on demand

    Captured 2026-08-08

    Arithmetic floor for 1 GPU over the contract's 730.5-hour month. This is not a provider offer.

    Open rental source

What to change

Computed actions and cost signals appear only when the result supports them.

RecommendationComputed memory fit

Consider FP8 or INT8 if the checkpoint and hardware support it

FP16 weights use 18.5 GB and bring total memory to 26.4 GB; FP8 or INT8 would lower weights to 9.2 GB and total memory to 13.7 GB, which fits inside 21.6 GB.

Show 2 more computed options
RecommendationComputed memory fit

Reduce context length or serving concurrency

Weights use 18.5 GB and fit inside 21.6 GB, but the 4.4 GB KV cache raises total memory to 26.4 GB; reduce --max-model-len or peak concurrency before adding hardware.

RecommendationComputed memory fit

Raise tensor parallel size to 2

The current vLLM command uses tensor parallel size 1 and provides 21.6 GB, while 26.4 GB requires 2 same-GPU workers for 43.2 GB; raise --tensor-parallel-size to 2 if the model topology supports it.

Decode bandwidth ceiling

A physics bound, not an expected throughput result.

32 output tok/sCeiling

CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.

Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters, as of 2026-08-08

Reproduce with vLLM

The exact command attached to this result.

vllm serve 'google/gemma-2-9b-it' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA A10 datasheet

    As of 2026-08-08

  • RunInfra Engine parameter-only KV-cache heuristic

    As of 2026-08-08

  • RunInfra Engine feasibility runtime-overhead policy

    As of 2026-08-09

  • Decode and prefill ceilings derived from GPU memory bandwidth, dense compute, and model parameters

    As of 2026-08-08

  • vLLM serve CLI documentation

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

  • NVIDIA H100 Tensor Core GPU datasheet

    As of 2026-08-08

Caveats

  • +Served context is 8192 tokens. It is a serving-capacity decision and is not inferred from average input or output tokens.
  • +Resolved layer count, KV-head count, or head dimension is absent. KV cache uses the parameter-only RunInfra Engine heuristic rather than an exact architecture calculation.
  • +Runtime overhead adds 3.44 GB to totalRequiredGb. This approximate 15% policy allowance covers CUDA context, activations, and vLLM overhead. Deployment-time profiling, including CUDA graph capture and allocator behavior, can differ.
  • +CEILING at 100% unit utilisation, not expected throughput. Decode is bounded by memory bandwidth and prefill by dense compute. Ideal tensor-parallel scaling is assumed for both. Real decode and prefill throughput are lower.
  • +Compute prefill ceiling: 6762.82064244471 tokens/sec. Bandwidth decode ceiling: 32.46153908373461 tokens/sec.
  • +CalcInput has no GPU provider selector. The calculation uses the lowest supplied on-demand rate for A10: OVHcloud, Public Cloud at $1/GPU-hour.
  • +The selected configuration does not fit, so no cost is returned.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Share this calculation

Embed code appears only when the calculator can publish a sourced cost for a fitting configuration.

Badge unavailable for this configuration

This configuration does not fit the selected hardware, so the calculator does not publish a cost or embed code.

Use it from an agent or script

GET the current inputs, or inspect the descriptor for accepted parameters and response fields.

Endpoint
GET https://runinfra.ai/api/calchttps://runinfra.ai/api/calc?hfId=google%2Fgemma-2-9b-it&gpuId=A10&gpuCount=1&requestsPerDay=50000&avgInputTokens=500&avgOutputTokens=500&utilization=0.3&quantization=fp16
Canonical result
https://runinfra.ai/calc/gemma-2-9b/a10/50k-req-day?context=8192
Descriptor
https://runinfra.ai/api/calc/schema

28 GPU ids, 7 quantizations

Make this configuration fit, describe what you need

Describe how you want to deploy this configuration...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy