RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/GPUs/B300
GPU

NVIDIA B300 for open model inference

We measured 1 published catalog package on this GPU. The rows below keep every claim tied to its recorded engine, concurrency, and verified date.

The one B300 package is also a serving recipe on unchanged weights, and the specification rows above stay unpublished until we can cite them.

The GPU specification is separate from our measured results

Memory, vendor spec
Memory not published.
Bandwidth, vendor spec
Bandwidth not published.
Hourly price, our rate
Hourly price not published.
Active price, our rate
Active hourly price not published.

The package rows show the measured differences

Measured package results on NVIDIA B300 with each row scoped to its published engine, concurrency, and verified date
ModelEngineTechniqueP50 latencyP99 latencyThroughputSpeed verdictAccuracyPrice and licenseConcurrencyVerified
Kimi K3
moonshotai/Kimi-K3
vLLM
0.23.1
Not published
19090 ms baseline
8978 ms optimized
19377 ms baseline
9595 ms optimized
55.2 tokens/s baseline
120 tokens/s optimized
2.12x faster
Parity by construction
$100 one-time
License Kimi K3 License
Concurrency 1
Jul 31, 2026

Published open-model inference measurements on NVIDIA B300, with baseline and optimized values only where the catalog carries both.

Measured by RunInfra on rented B300 hardware with vLLM 0.23.1; signed benchmark receipt ships in each package kit.

We report package measurements here, and each package remains subject to its listed license.

Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.

Throughput changes by model and package

Published throughput pairs already shown in the package table
Kimi K3
Baseline
55.2 tokens/s
Optimized
120 tokens/s

What is not measured

Other GPU measurements for these listed packages not published.

Other engine measurements for these listed packages not published.

Measurements for concurrency not listed in the table for these packages not published.

The same measurements answer common questions

What p50 latency pair is published for Kimi K3 on NVIDIA B300?

Kimi K3 p50 latency measured 19090 ms at baseline and 8978 ms optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.

What p99 latency pair is published for Kimi K3 on NVIDIA B300?

Kimi K3 p99 latency measured 19377 ms at baseline and 9595 ms optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.

What throughput pair is published for Kimi K3 on NVIDIA B300?

Kimi K3 throughput measured 55.2 tokens/s at baseline and 120 tokens/s optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy