RunInfraby RightNow
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

See what this model actually costs to serve.

Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.

Qwen2.5-32B-Instruct fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established.

Fit is always evaluated when model sizing is available. GPU rent, paid idle capacity, API crossover, and measured optimization appear only when their required evidence exists. When the API wins, the page says so.

Model

Qwen/Qwen2.5-32B-Instruct

GPU

NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB

Requests/day

50,000

GPU count

1

Quantization

FP16

Tokens/request

500 in, 500 out

Charts appear only when evidence existsUtilization defaults to 30%

Configure the workload

Change the inputs, then calculate a shareable result.

Hugging Face model

Utilization assumption

Defaulted to 30% so the page does not manufacture a self-hosting win.

Derived utilization

Needs throughput

Idle share paid

Needs throughput

Utilization assumption30%

A steady production service with real peaks and troughs.

At this level about 70% of every rented GPU hour sits idle, and it is billed at the same rate as a busy one.

Estimated memory fit

This configuration is not runnable as entered.

No complete throughput evidence can price this configuration. Memory fit and provenance remain available while money and comparisons stay absent.

Model

Qwen/Qwen2.5-32B-Instruct

Hardware

1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, FP16

Memory fit

Weights and KV cache stay separate so the fit is reproducible.

65.5 GB
Model weights
2.1 GB
KV cache
77.8 GB
Total required
86.4 GB
Available
8,192 tokens
Served context used
10.2 GB
Runtime overhead
Exact transformer shape
KV cache method

Decode bandwidth ceiling

A physics bound, not an expected throughput result.

24 output tok/sCeiling

Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.

Decode ceiling derived from GPU memory bandwidth and model weight bytes, as of 2026-08-08

Provider price spread, NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB

Dated GPU-hour rental prices for this hardware. They are not monthly serving totals because throughput evidence is missing.

Scroll horizontally to read every chart label and value.

AWS, US East (N. Virginia): $3.3631, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Google Cloud, Iowa us-central1: $4.4999, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Cheapest: Hyperstack: $1.85, on demand GPU-hour rate, captured 2026-08-08 (1 day ago). Microsoft Azure, East US 2: $5.50, on demand GPU-hour rate, captured 2026-08-08 (1 day ago).

  • AWS, US East (N. Virginia)$3.3631
  • Google Cloud, Iowa us-central1$4.4999
  • Cheapest: Hyperstack$1.85
  • Microsoft Azure, East US 2$5.50
Source:
  • AWS, US East (N. Virginia), on demand, 1 day ago, captured 2026-08-08
  • Google Cloud, Iowa us-central1, on demand, 1 day ago, captured 2026-08-08
  • Hyperstack, on demand, 1 day ago, captured 2026-08-08
  • Microsoft Azure, East US 2, on demand, 1 day ago, captured 2026-08-08
+What this chart means

These are GPU-hour rental prices only. A monthly serving total requires throughput evidence and remains absent.

vLLM deployment starting point

This line starts the selected vLLM shape. It is not a throughput reproduction claim.

vllm serve 'Qwen/Qwen2.5-32B-Instruct' --dtype float16 --tensor-parallel-size 1 --max-model-len 8192 --gpu-memory-utilization 0.9

Not measured on this config yet

No site-wide multiplier is applied. Run the exact configuration to create a receipt before comparing its measured baseline with RunInfra.

Run this configuration

Sources and caveats

Every published number keeps its source and date.

Provenance

  • Hugging Face model API revision sha

    As of 2026-08-08

  • Hugging Face model API safetensors.total

    As of 2026-08-08

  • config.json:max_position_embeddings

    As of 2026-08-08

  • config.json:architectures[0]

    As of 2026-08-08

  • NVIDIA RTX PRO AI Factory, RTX PRO 6000 Blackwell Server Edition specifications

    As of 2026-08-08

  • Hugging Face config.json transformer shape

    As of 2026-08-08

  • RunInfra Engine feasibility runtime-overhead policy

    As of 2026-08-09

  • Decode ceiling derived from GPU memory bandwidth and model weight bytes

    As of 2026-08-08

  • vLLM serve CLI documentation

    As of 2026-08-08

  • NVIDIA H200 Tensor Core GPU datasheet

    As of 2026-08-08

  • NVIDIA Blackwell architecture specification

    As of 2026-08-08

  • NVIDIA Blackwell Ultra Datasheet, page 5, HGX B300 individual-GPU sparse specifications; footnote 2: 'Specification in sparse. Dense is ½ sparse spec shown.' Dense fp16 is 4.5 / 2 = 2.25 PFLOPS and dense fp8 is 9 / 2 = 4.5 PFLOPS

    As of 2026-08-09

  • NVIDIA H100 NVL product brief

    As of 2026-08-08

Caveats

  • +Served context is 8192 tokens. It is a serving-capacity decision and is not inferred from average input or output tokens.
  • +KV cache uses the exact transformer-shape formula with 64 layers, 8 KV heads, and head dimension 128.
  • +Runtime overhead adds 10.15 GB to totalRequiredGb. This approximate 15% policy allowance covers CUDA context, activations, and vLLM overhead. Deployment-time profiling, including CUDA graph capture and allocator behavior, can differ.
  • +Bandwidth decode CEILING, not expected throughput. Ideal tensor-parallel scaling is assumed. Real decode throughput is lower.
  • +No dense compute ceiling is cited for fp16 on NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB. The available physics bound is decode-only, so the entry point returns no cost instead of copying decode throughput into the input field.
  • +CalcInput has no GPU provider selector. The calculation uses the lowest supplied on-demand rate for RTX-PRO-6000-Blackwell-Server: Hyperstack at $1.85/GPU-hour.

Share this calculation

Embed code appears only when the calculator can publish a sourced cost for a fitting configuration.

Badge unavailable for this configuration

The calculator does not have a complete throughput source for this workload, so it withheld the cost and does not generate embed code.

Use it from an agent or script

GET the current inputs, or inspect the descriptor for accepted parameters and response fields.

Endpoint
GET https://runinfra.ai/api/calchttps://runinfra.ai/api/calc?hfId=Qwen%2FQwen2.5-32B-Instruct&gpuId=RTX-PRO-6000-Blackwell-Server&gpuCount=1&requestsPerDay=50000&avgInputTokens=500&avgOutputTokens=500&utilization=0.3&quantization=fp16&servedContextTokens=8192
Canonical result
https://runinfra.ai/calc/qwen-2.5-32b/rtx-pro-6000-blackwell-server/50k-req-day
Descriptor
https://runinfra.ai/api/calc/schema

28 GPU ids, 7 quantizations

Turn this verified configuration into a deployment, describe what you need

Describe how you want to deploy this configuration...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy