What p50 latency pair is published for Kimi K3 on NVIDIA B300?
Kimi K3 p50 latency measured 19090 ms at baseline and 8978 ms optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.
We measured 1 published catalog package on this GPU. The rows below keep every claim tied to its recorded engine, concurrency, and verified date.
The one B300 package is also a serving recipe on unchanged weights, and the specification rows above stay unpublished until we can cite them.
| Model | Engine | Technique | P50 latency | P99 latency | Throughput | Speed verdict | Accuracy | Price and license | Concurrency | Verified |
|---|---|---|---|---|---|---|---|---|---|---|
| Kimi K3 moonshotai/Kimi-K3 | vLLM 0.23.1 | Not published | 19090 ms baseline 8978 ms optimized | 19377 ms baseline 9595 ms optimized | 55.2 tokens/s baseline 120 tokens/s optimized | 2.12x faster | Parity by construction | $100 one-time License Kimi K3 License | Concurrency 1 | Jul 31, 2026 |
Published open-model inference measurements on NVIDIA B300, with baseline and optimized values only where the catalog carries both.
Measured by RunInfra on rented B300 hardware with vLLM 0.23.1; signed benchmark receipt ships in each package kit.
We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.
Other GPU measurements for these listed packages not published.
Other engine measurements for these listed packages not published.
Measurements for concurrency not listed in the table for these packages not published.
Kimi K3 p50 latency measured 19090 ms at baseline and 8978 ms optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.
Kimi K3 p99 latency measured 19377 ms at baseline and 9595 ms optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.
Kimi K3 throughput measured 55.2 tokens/s at baseline and 120 tokens/s optimized under concurrency 1, vLLM 0.23.1, measurement verified Jul 31, 2026. Package verdict: 2.12x faster.
© 2026 RunInfra. All rights reserved.