RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/GPUs/H100
GPU

NVIDIA H100 for open model inference

We measured 3 published catalog packages on this GPU. The rows below keep every claim tied to its recorded engine, concurrency, and verified date.

Each of the three H100 packages takes the same lever, channelwise FP8, and each keeps its measured accuracy within a fraction of a percent of its BF16 baseline.

The GPU specification is separate from our measured results

Memory, vendor spec
80 GB
Bandwidth, vendor spec
3,350 GB/s
Hourly price, our rate
$4.18 per GPU-hour
Active price, our rate
$3.35 per GPU-hour

The package rows show the measured differences

Measured package results on NVIDIA H100 with each row scoped to its published engine, concurrency, and verified date
ModelEngineTechniqueP50 latencyP99 latencyThroughputSpeed verdictAccuracyPrice and licenseConcurrencyVerified
AREX-Turbo
BAAI/AREX-Turbo
vLLM
0.25.1
Channelwise FP8 weights, dynamic per-token activations
666 ms baseline
551 ms optimized
684 ms baseline
579 ms optimized
1534 tokens/s baseline
1847 tokens/s optimized
1.2x faster
gsm8k no measurable accuracy change, passed
gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2): 0.3821 baseline to 0.4094 optimized
$15 one-time
License Apache-2.0
Concurrency 8
Jul 27, 2026
Qwen3.6 27B
Qwen/Qwen3.6-27B
vLLM
0.25.1
Channelwise FP8, measured selective-layer recipe
2857 ms baseline
2215 ms optimized
2881 ms baseline
2230 ms optimized
361 tokens/s baseline
465 tokens/s optimized
1.28x faster
gsm8k 99.87% recovery, passed
gsm8k exact_match strict (N=1319, completion protocol): 0.5754 baseline to 0.5747 optimized
$40 one-time
License Apache-2.0
Concurrency 8
Jul 25, 2026
Qwythos-9B-Claude-Mythos-5-1M
empero-ai/Qwythos-9B-Claude-Mythos-5-1M
vLLM
0.25.1
Channelwise FP8 weights, dynamic per-token activations
1058 ms baseline
815 ms optimized
1083 ms baseline
861 ms optimized
974 tokens/s baseline
1255 tokens/s optimized
1.29x faster
gsm8k 99.35% recovery, passed
gsm8k exact_match strict (N=1319, completion protocol): 0.8378 baseline to 0.8324 optimized
$20 one-time
License Apache-2.0
Concurrency 8
Jul 25, 2026

Published open-model inference measurements on NVIDIA H100, with baseline and optimized values only where the catalog carries both.

Measured by RunInfra on rented H100 hardware with vLLM 0.25.1; signed benchmark receipt ships in each package kit.

We report package measurements here, and each package remains subject to its listed license.

Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.

Throughput changes by model and package

Published throughput pairs already shown in the package table
AREX-Turbo
Baseline
1534 tokens/s
Optimized
1847 tokens/s
Qwen3.6 27B
Baseline
361 tokens/s
Optimized
465 tokens/s
Qwythos-9B-Claude-Mythos-5-1M
Baseline
974 tokens/s
Optimized
1255 tokens/s

What is not measured

Other GPU measurements for these listed packages not published.

Other engine measurements for these listed packages not published.

Measurements for concurrency not listed in the table for these packages not published.

The same measurements answer common questions

What p50 latency pair is published for AREX-Turbo on NVIDIA H100?

AREX-Turbo p50 latency measured 666 ms at baseline and 551 ms optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 27, 2026. Package verdict: 1.2x faster.

What p50 latency pair is published for Qwen3.6 27B on NVIDIA H100?

Qwen3.6 27B p50 latency measured 2857 ms at baseline and 2215 ms optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 25, 2026. Package verdict: 1.28x faster.

What p50 latency pair is published for Qwythos-9B-Claude-Mythos-5-1M on NVIDIA H100?

Qwythos-9B-Claude-Mythos-5-1M p50 latency measured 1058 ms at baseline and 815 ms optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 25, 2026. Package verdict: 1.29x faster.

What p99 latency pair is published for AREX-Turbo on NVIDIA H100?

AREX-Turbo p99 latency measured 684 ms at baseline and 579 ms optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 27, 2026. Package verdict: 1.2x faster.

What throughput pair is published for AREX-Turbo on NVIDIA H100?

AREX-Turbo throughput measured 1534 tokens/s at baseline and 1847 tokens/s optimized under concurrency 8, vLLM 0.25.1, measurement verified Jul 27, 2026. Package verdict: 1.2x faster.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy