RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/GPUs/B200
GPU

NVIDIA B200 for open model inference

We measured 1 published catalog package on this GPU. The rows below keep every claim tied to its recorded engine, concurrency, and verified date.

The B200 gain is a serving-stack result on unchanged weights, speculative decoding the released engine does not enable on its own, so accuracy holds by construction.

The GPU specification is separate from our measured results

Memory, vendor spec
192 GB
Bandwidth, vendor spec
8,000 GB/s
Hourly price, our rate
$8.64 per GPU-hour
Active price, our rate
$6.84 per GPU-hour

The package rows show the measured differences

Measured package results on NVIDIA B200 with each row scoped to its published engine, concurrency, and verified date
ModelEngineTechniqueP50 latencyP99 latencyThroughputSpeed verdictAccuracyPrice and licenseConcurrencyVerified
DeepSeek V4 Flash
deepseek-ai/DeepSeek-V4-Flash-0731
vLLM
0.25.0
Not published
9248 ms baseline
3095 ms optimized
9548 ms baseline
4306 ms optimized
113 tokens/s baseline
363 tokens/s optimized
2.98x faster
Parity by construction
$60 one-time
License MIT
Concurrency 1
Aug 1, 2026

Published open-model inference measurements on NVIDIA B200, with baseline and optimized values only where the catalog carries both.

Measured by RunInfra on rented B200 hardware with vLLM 0.25.0; signed benchmark receipt ships in each package kit.

We report package measurements here, and each package remains subject to its listed license.

Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.

Throughput changes by model and package

Published throughput pairs already shown in the package table
DeepSeek V4 Flash
Baseline
113 tokens/s
Optimized
363 tokens/s

What is not measured

Other GPU measurements for these listed packages not published.

Other engine measurements for these listed packages not published.

Measurements for concurrency not listed in the table for these packages not published.

The same measurements answer common questions

What p50 latency pair is published for DeepSeek V4 Flash on NVIDIA B200?

DeepSeek V4 Flash p50 latency measured 9248 ms at baseline and 3095 ms optimized under concurrency 1, vLLM 0.25.0, measurement verified Aug 1, 2026. Package verdict: 2.98x faster.

What p99 latency pair is published for DeepSeek V4 Flash on NVIDIA B200?

DeepSeek V4 Flash p99 latency measured 9548 ms at baseline and 4306 ms optimized under concurrency 1, vLLM 0.25.0, measurement verified Aug 1, 2026. Package verdict: 2.98x faster.

What throughput pair is published for DeepSeek V4 Flash on NVIDIA B200?

DeepSeek V4 Flash throughput measured 113 tokens/s at baseline and 363 tokens/s optimized under concurrency 1, vLLM 0.25.0, measurement verified Aug 1, 2026. Package verdict: 2.98x faster.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy