RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Catalog/Kimi K3
Model

Kimi K3 measured packages and prices

RunInfra has published 1 measured package for moonshotai/Kimi-K3 on B300. Every package below keeps its engine, price, and proof date together.

Parameters
2800B
Modality
Text
Base model license
Kimi K3 License
Hugging Face model
moonshotai/Kimi-K3

The published packages are available to buy

Published packages for Kimi K3, with GPU, engine, optimization, price, and proof date
PackageGPUEngineOptimizationPriceVerified
kimi-k3-b300x8-k3turboB300vLLM 0.23.1Optimization technique not published.$100 one-timeJul 31, 2026

What we measured

The package keeps the upstream weights unchanged and sells the validated serving profiles, build recipe, and proof instead.

kimi-k3-b300x8-k3turbo

Published Jul 31, 2026. Proof verified Jul 31, 2026.

2.12x faster

Median latency measured 19090 ms at baseline and 8978 ms optimized.

Quality evidence

Parity by construction.

Published package measurements and prices for Kimi K3, grouped from the public RunInfra catalog.

Measured by RunInfra on B300 with vLLM 0.23.1; proof dates are listed with each package.

We report package measurements here, and each package remains subject to its listed license.

Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.

What is not measured

Other GPU measurements for this model not published.

Other engine measurements for this model not published.

Measurements for context lengths not listed in these packages not published.

Published facts answer common questions

What does Kimi K3 cost on RunInfra?

kimi-k3-b300x8-k3turbo is listed at $100 one-time, published Jul 31, 2026.

Which GPU serves Kimi K3?

kimi-k3-b300x8-k3turbo is measured on B300 with vLLM 0.23.1. Proof verified Jul 31, 2026.

Is quality measured for Kimi K3?

Parity by construction. Proof verified Jul 31, 2026.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy