RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

See what this model actually costs to serve.

Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.

See what this model actually costs to serve.

Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.

Model

black-forest-labs/FLUX.1-schnell

GPU

NVIDIA H100 SXM 80 GB

Requests/day

50,000

GPU count

1

Quantization

FP16

Tokens/request

500 in, 500 out

Charts appear only when evidence existsUtilization defaults to 30%

Configure the workload

Change the inputs, then calculate a shareable result.

Hugging Face model

Utilization assumption

Defaulted to 30% so the page does not manufacture a self-hosting win.

Derived utilization

Needs throughput

Idle share paid

Needs throughput

Utilization assumption30%

A steady production service with real peaks and troughs.

At this level about 70% of every rented GPU hour sits idle, and it is billed at the same rate as a busy one.

Fit confirmed

Cost needs a throughput source.

The fit verdict, compatible GPUs, and reproducing command are still useful. Money stays absent until this traffic shape has a complete measured, cited, or physics-bound evidence path.

Model

black-forest-labs/FLUX.1-schnell

Hardware

1 x NVIDIA H100 SXM 80 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

23.8 GB
Model weights
0.7 GB
KV cache
24.5 GB
Total required
72 GB
Available

Provider price spread, NVIDIA H100 SXM 80 GB

Dated GPU-hour rental prices for this hardware. They are not monthly serving totals because throughput evidence is missing.

Scroll horizontally to read every chart label and value.

DigitalOcean, GPU Droplets: $4.41, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Hyperstack: $3.20, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Lambda Cloud: $4.29, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Nebius: $3.85, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest: RunPod, Community Cloud: $2.69, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). RunPod, Secure Cloud: $2.99, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Verda: $3.25, on demand GPU-hour rate, captured 2026-08-08 (1 day ago).

  • DigitalOcean, GPU Droplets$4.41
  • Hyperstack$3.20
  • Lambda Cloud$4.29
  • Nebius$3.85
  • Cheapest: RunPod, Community Cloud$2.69
  • RunPod, Secure Cloud$2.99
  • Verda$3.25
Source:
  • DigitalOcean, GPU Droplets, on demand, 1 day ago, captured 2026-08-08
  • Hyperstack, on demand, 1 day ago, captured 2026-08-08
  • Lambda Cloud, on demand, 1 day ago, captured 2026-08-08
  • Nebius, on demand, 1 day ago, captured 2026-08-08
  • RunPod, Community Cloud, on demand, 1 day ago, captured 2026-08-08
  • RunPod, Secure Cloud, on demand, 1 day ago, captured 2026-08-08
  • Verda, on demand, 1 day ago, captured 2026-08-08
+What this chart means

These are GPU-hour rental prices only. A monthly serving total requires throughput evidence and remains absent.

Not measured on this config yet

No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.

Run this configuration

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA H100 Tensor Core GPU datasheet

    As of 2026-08-08

  • RunInfra Engine parameter-only KV-cache heuristic

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

Caveats

  • +Workload class diffusion has no class-specific activation shape in CalcDataset. The kvCacheGb field is a parameter-based runtime-memory allowance used only for the VRAM budget. It is not an autoregressive KV-cache measurement, a backend compatibility test, or a serving performance estimate.
  • +Class-specific activation dimensions are absent. Fit uses the configured input plus output token count only to scale its parameter-based runtime-memory allowance; that token shape is not presented as a workload-native request unit.
  • +Workload class diffusion is not a decoder language model covered by this calculator. Memory fit remains available when model size is known, but output-token throughput, serving cost, the cost-vs-volume curve, API comparisons, the optimized delta, and a vLLM serve command are withheld because this workload's serving path and economic unit are not modelled here.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Fit confirmed

Cost needs a throughput source.

The fit verdict, compatible GPUs, and reproducing command are still useful. Money stays absent until this traffic shape has a complete measured, cited, or physics-bound evidence path.

Model

black-forest-labs/FLUX.1-schnell

Hardware

1 x NVIDIA H100 SXM 80 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

23.8 GB
Model weights
0.7 GB
KV cache
24.5 GB
Total required
72 GB
Available

Provider price spread, NVIDIA H100 SXM 80 GB

Dated GPU-hour rental prices for this hardware. They are not monthly serving totals because throughput evidence is missing.

Scroll horizontally to read every chart label and value.

DigitalOcean, GPU Droplets: $4.41, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Hyperstack: $3.20, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Lambda Cloud: $4.29, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Nebius: $3.85, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest: RunPod, Community Cloud: $2.69, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). RunPod, Secure Cloud: $2.99, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Verda: $3.25, on demand GPU-hour rate, captured 2026-08-08 (1 day ago).

  • DigitalOcean, GPU Droplets$4.41
  • Hyperstack$3.20
  • Lambda Cloud$4.29
  • Nebius$3.85
  • Cheapest: RunPod, Community Cloud$2.69
  • RunPod, Secure Cloud$2.99
  • Verda$3.25
Source:
  • DigitalOcean, GPU Droplets, on demand, 1 day ago, captured 2026-08-08
  • Hyperstack, on demand, 1 day ago, captured 2026-08-08
  • Lambda Cloud, on demand, 1 day ago, captured 2026-08-08
  • Nebius, on demand, 1 day ago, captured 2026-08-08
  • RunPod, Community Cloud, on demand, 1 day ago, captured 2026-08-08
  • RunPod, Secure Cloud, on demand, 1 day ago, captured 2026-08-08
  • Verda, on demand, 1 day ago, captured 2026-08-08
+What this chart means

These are GPU-hour rental prices only. A monthly serving total requires throughput evidence and remains absent.

Not measured on this config yet

No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.

Run this configuration

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA H100 Tensor Core GPU datasheet

    As of 2026-08-08

  • RunInfra Engine parameter-only KV-cache heuristic

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

Caveats

  • +Workload class diffusion has no class-specific activation shape in CalcDataset. The kvCacheGb field is a parameter-based runtime-memory allowance used only for the VRAM budget. It is not an autoregressive KV-cache measurement, a backend compatibility test, or a serving performance estimate.
  • +Class-specific activation dimensions are absent. Fit uses the configured input plus output token count only to scale its parameter-based runtime-memory allowance; that token shape is not presented as a workload-native request unit.
  • +Workload class diffusion is not a decoder language model covered by this calculator. Memory fit remains available when model size is known, but output-token throughput, serving cost, the cost-vs-volume curve, API comparisons, the optimized delta, and a vLLM serve command are withheld because this workload's serving path and economic unit are not modelled here.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Fit confirmed

Cost needs a throughput source.

The fit verdict, compatible GPUs, and reproducing command are still useful. Money stays absent until this traffic shape has a complete measured, cited, or physics-bound evidence path.

Model

black-forest-labs/FLUX.1-schnell

Hardware

1 x NVIDIA H100 SXM 80 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

23.8 GB
Model weights
0.7 GB
KV cache
24.5 GB
Total required
72 GB
Available

Provider price spread, NVIDIA H100 SXM 80 GB

Dated GPU-hour rental prices for this hardware. They are not monthly serving totals because throughput evidence is missing.

Scroll horizontally to read every chart label and value.

DigitalOcean, GPU Droplets: $4.41, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Hyperstack: $3.20, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Lambda Cloud: $4.29, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Nebius: $3.85, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest: RunPod, Community Cloud: $2.69, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). RunPod, Secure Cloud: $2.99, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Verda: $3.25, on demand GPU-hour rate, captured 2026-08-08 (1 day ago).

  • DigitalOcean, GPU Droplets$4.41
  • Hyperstack$3.20
  • Lambda Cloud$4.29
  • Nebius$3.85
  • Cheapest: RunPod, Community Cloud$2.69
  • RunPod, Secure Cloud$2.99
  • Verda$3.25
Source:
  • DigitalOcean, GPU Droplets, on demand, 1 day ago, captured 2026-08-08
  • Hyperstack, on demand, 1 day ago, captured 2026-08-08
  • Lambda Cloud, on demand, 1 day ago, captured 2026-08-08
  • Nebius, on demand, 1 day ago, captured 2026-08-08
  • RunPod, Community Cloud, on demand, 1 day ago, captured 2026-08-08
  • RunPod, Secure Cloud, on demand, 1 day ago, captured 2026-08-08
  • Verda, on demand, 1 day ago, captured 2026-08-08
+What this chart means

These are GPU-hour rental prices only. A monthly serving total requires throughput evidence and remains absent.

Not measured on this config yet

No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.

Run this configuration

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA H100 Tensor Core GPU datasheet

    As of 2026-08-08

  • RunInfra Engine parameter-only KV-cache heuristic

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

Caveats

  • +Workload class diffusion has no class-specific activation shape in CalcDataset. The kvCacheGb field is a parameter-based runtime-memory allowance used only for the VRAM budget. It is not an autoregressive KV-cache measurement, a backend compatibility test, or a serving performance estimate.
  • +Class-specific activation dimensions are absent. Fit uses the configured input plus output token count only to scale its parameter-based runtime-memory allowance; that token shape is not presented as a workload-native request unit.
  • +Workload class diffusion is not a decoder language model covered by this calculator. Memory fit remains available when model size is known, but output-token throughput, serving cost, the cost-vs-volume curve, API comparisons, the optimized delta, and a vLLM serve command are withheld because this workload's serving path and economic unit are not modelled here.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Fit confirmed

Cost needs a throughput source.

The fit verdict, compatible GPUs, and reproducing command are still useful. Money stays absent until this traffic shape has a complete measured, cited, or physics-bound evidence path.

Model

black-forest-labs/FLUX.1-schnell

Hardware

1 x NVIDIA H100 SXM 80 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

23.8 GB
Model weights
0.7 GB
KV cache
24.5 GB
Total required
72 GB
Available

Provider price spread, NVIDIA H100 SXM 80 GB

Dated GPU-hour rental prices for this hardware. They are not monthly serving totals because throughput evidence is missing.

Scroll horizontally to read every chart label and value.

DigitalOcean, GPU Droplets: $4.41, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Hyperstack: $3.20, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Lambda Cloud: $4.29, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Nebius: $3.85, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest: RunPod, Community Cloud: $2.69, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). RunPod, Secure Cloud: $2.99, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Verda: $3.25, on demand GPU-hour rate, captured 2026-08-08 (1 day ago).

  • DigitalOcean, GPU Droplets$4.41
  • Hyperstack$3.20
  • Lambda Cloud$4.29
  • Nebius$3.85
  • Cheapest: RunPod, Community Cloud$2.69
  • RunPod, Secure Cloud$2.99
  • Verda$3.25
Source:
  • DigitalOcean, GPU Droplets, on demand, 1 day ago, captured 2026-08-08
  • Hyperstack, on demand, 1 day ago, captured 2026-08-08
  • Lambda Cloud, on demand, 1 day ago, captured 2026-08-08
  • Nebius, on demand, 1 day ago, captured 2026-08-08
  • RunPod, Community Cloud, on demand, 1 day ago, captured 2026-08-08
  • RunPod, Secure Cloud, on demand, 1 day ago, captured 2026-08-08
  • Verda, on demand, 1 day ago, captured 2026-08-08
+What this chart means

These are GPU-hour rental prices only. A monthly serving total requires throughput evidence and remains absent.

Not measured on this config yet

No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.

Run this configuration

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • Hugging Face config.json

    As of 2026-08-08

  • NVIDIA H100 Tensor Core GPU datasheet

    As of 2026-08-08

  • RunInfra Engine parameter-only KV-cache heuristic

    As of 2026-08-08

  • NVIDIA L40S product specification

    As of 2026-08-08

  • NVIDIA A100 product specification

    As of 2026-08-08

Caveats

  • +Workload class diffusion has no class-specific activation shape in CalcDataset. The kvCacheGb field is a parameter-based runtime-memory allowance used only for the VRAM budget. It is not an autoregressive KV-cache measurement, a backend compatibility test, or a serving performance estimate.
  • +Class-specific activation dimensions are absent. Fit uses the configured input plus output token count only to scale its parameter-based runtime-memory allowance; that token shape is not presented as a workload-native request unit.
  • +Workload class diffusion is not a decoder language model covered by this calculator. Memory fit remains available when model size is known, but output-token throughput, serving cost, the cost-vs-volume curve, API comparisons, the optimized delta, and a vLLM serve command are withheld because this workload's serving path and economic unit are not modelled here.
  • +The Hugging Face repository is gated. Running the generated command requires authorized access to the model weights.

Share this calculation

Embed code appears only when the calculator can publish a sourced cost for a fitting configuration.

Badge unavailable for this configuration

The calculator does not have a complete throughput source for this workload, so it withheld the cost and does not generate embed code.

Use it from an agent or script

GET the current inputs, or inspect the descriptor for accepted parameters and response fields.

Endpoint
GET https://runinfra.ai/api/calchttps://runinfra.ai/api/calc?hfId=black-forest-labs%2FFLUX.1-schnell&gpuId=H100&gpuCount=1&requestsPerDay=50000&avgInputTokens=500&avgOutputTokens=500&utilization=0.3&quantization=fp16
Canonical result
https://runinfra.ai/calc/flux-1-schnell/h100/50k-req-day
Descriptor
https://runinfra.ai/api/calc/schema

28 GPU ids, 7 quantizations

Turn this exact configuration into a deployment, describe what you need

Describe how you want to deploy this configuration...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy