Serving-tuned for B300 on vLLM vLLM 0.23.1rc1.dev1475+gf68f4fdde (k3-turbo wheel, vllm@f68f4fdd), with the benchmark receipt included. You get the engineering: the exact serving configuration, the measured proof, and the verifier. The weights come from Hugging Face at a pinned revision, about 1561.0 GB, so we redistribute nothing and the bytes you run are the bytes we measured. Run it on any cloud or your own hardware. No RunInfra service is required at runtime.
Weights are not in the kit
This package is a recipe. The kit carries the serving configuration, the benchmark receipt and the verifier. The weights come from Hugging Face at a pinned revision, so nothing is redistributed and the bytes you run are the bytes we measured.
We keep re-optimizing and re-benchmarking as engines move, and you always download the latest verified version. Cancel anytime. Nothing is charged until you confirm at checkout.
On. Confirmed at checkout.
Latency p50
18695ms, concurrency 8
Latency p99
18766ms, concurrency 8
Accuracy
Served-output parity with the pinned upstream engine build (by construction; no measured evaluation)Accuracy recovery not measuredPassed accuracy gate
GPU target
B300
Base model license
Kimi K3 License
Proof verified
Jul 29, 2026
Published
Jul 29, 2026
Runs on any GPU provider
Bring the kit to any cloud with B300-class GPUs, or your own hardware, and pull the weights from Hugging Face at the pinned revision.
Docker Compose
This is what it means to own your AI
You control the intelligence
The weights run on your infrastructure. No deprecation notices, no silent model swaps, no provider deciding what your product is allowed to do.
You keep the economics
One purchase instead of a meter that runs forever. A closed API charges you for every token, at whatever the price becomes; your own deployment costs what your GPU costs.
You keep the data
Prompts and outputs never leave your network. Inference runs on your infrastructure without calling RunInfra, while the weights fetch comes from Hugging Face at the pinned revision.
Two licenses apply: the RunInfra package license you purchase under, and the base model's own open-source license.
Package license
RunInfra Package License v1.1 (2026-07-25)
Buy once. No meter and no expiry. The model weights come from Hugging Face at the pinned revision, and no RunInfra service is required at runtime. The full terms ship inside the kit.
What you are licensed to do
This package is licensed to the purchasing workspace, one time, for the exact version purchased. Everyone in that workspace may run it in production without limits: unlimited inference, on any hardware or cloud the workspace controls, for any lawful commercial purpose. You may modify the configuration and tooling for your own use.
Your outputs are yours
Everything the model produces for you belongs to you. You may use, sell, and build products on the model's outputs without restriction or royalty. Serving the model to your own customers as part of your product is use, not redistribution, and is fully allowed.
What you may not do
You may not redistribute, resell, sublicense, rent, publish, or otherwise make the package or its artifacts available to any third party. That covers the serving configuration, scripts, kit archive, and benchmark receipts, whole or in part, modified or not. The upstream model weights are not included in this kit. You fetch them separately from Hugging Face at the pinned revision, and they remain governed by the model's own license, not this package license. The license belongs to the purchasing workspace and cannot be transferred separately from it.
The exact engine build and serve flags proven healthy on 8x B300 on 2026-07-29, written flag for flag into every deploy file: the pinned source-built vLLM wheel (vllm at commit f68f4fdd), tensor parallel 8 with expert parallelism, 40960 context, the Marlin MXFP4 mixture-of-experts backend for K3's SiTU activation, the Triton KDA prefill backend, NCCL all-reduce in place of the custom all-reduce that fails at CUDA graph capture on this node, 128-token KV blocks for FlashInfer MLA alignment, and the expandable-segments allocator. No stock configuration exists to compare against, so this configuration is sold as what it is, the validated way to serve the model, and it is never counted into a speedup multiplier.
Measured accuracy
Served-output parity with the pinned upstream engine build
(by construction; no measured evaluation)
Passed accuracy gate
Baseline
Not reported
Optimized
Not reported
Recovery
Not reported
Delta
Not reported
Accuracy recovery not measured
This package does not publish both a baseline and an optimized score on this metric, so there is no recovery figure to state.
Methodology
By construction, not a measured evaluation: v1 changes no weight byte (the buyer-fetched tree is verified at deploy time against the measured digest ebf0b1a75a29ab5e01f319252302c418d0e8f9340e402950497094a6fc64aced) and applies no kernel or runtime modification (recipe kernel_stages_applied is empty and the engagement probe proved the plugin leaves every serving symbol stock), so the served model is the pinned upstream vLLM build serving the pinned upstream MXFP4 weights and its outputs are that build's own outputs.
No cross-server token-agreement score is published because the method is unsound here: a stock-vs-stock control between two identical server instances measured 78.12 percent top-1 agreement (results/raw/trial-payload.json), so cross-server logit comparison at tensor parallel 8 cannot certify accuracy and is informational only.
Runtime not included in the kit
Optimized measurement: These numbers were measured on a source-built engine wheel, vLLM 0.23.1rc1.dev1475+gf68f4fdde, compiled from vllm-project/vllm at commit f68f4fdd (the PR 50000 head) with the in-tree FlashKDA CUDA extension disabled, built and run on Modal. The kit does not ship that wheel binary. It ships the pinned build recipe for it (deploy/build/Dockerfile.vllm-k3-wheel, same commit, same exclusions), which had not yet been executed on a Docker-capable machine at verification time.
Buyer disclosure
What we did not measure
Refusal, toxicity and jailbreak behavior were not evaluated. This package does not modify the weights in any way, so we do not expect it to shift them, but we did not measure it. Run your own safety evaluation before user-facing or regulated deployment.
Kernel-level speedup claims are NOT made for Kimi K3 in this version and none appear in the numbers: the day-0 upstream vLLM tree ships K3 already fused where the k3-turbo kernel set operates (in-tree fused decode is a superset of the S3 kernel; the prefill chunk chain is 2 to 3 percent of time to first token, below the 2 percent end-to-end improvement contract). The S3, F7 and F13 kernels in this kit are measured wins on the Kimi-Linear model class and are included with their proxy-model receipts. This package's measured Kimi K3 value is the validated serving configuration itself.
No baseline row is published, and the absence is the finding: as of 2026-07-29 stock vLLM cannot serve Kimi K3 through any published wheel or public default path. No released vLLM wheel carries K3 support, public flashinfer wheels lack the SiTU activation type K3's mixture-of-experts layers require, FlashKDA prefill depends on a private CUDA extension, custom all-reduce failed at the tail of full CUDA graph capture on this node, and FlashInfer MLA imposes a KV block alignment the defaults do not satisfy. This package exists because it resolves those five walls with a pinned source build and a proven flag set, found across 15 debugging cycles. There is no stock configuration to measure a speedup against, so no speedup is claimed: the numbers published are absolute measurements of this configuration, which is, to our knowledge, the first publicly-reproducible way to serve this model.
The underlying model remains under its upstream open-source license, which is included in this kit with attribution and a statement of RunInfra's modifications. Nothing in this license restricts rights the upstream license grants you for the ORIGINAL model; the restrictions above apply to RunInfra's optimized package.
Version-pinned, as measured
You purchased this exact version, proven against the exact engine version named in the receipt. It stays downloadable to your workspace and never expires. The benchmark receipt describes measurements taken at verification time on the named hardware; RunInfra does not promise future updates to this version, and later package versions are separate purchases.
Breach ends the license
If the workspace redistributes the package or its artifacts, this license terminates for that workspace. Sections about your outputs survive termination for outputs already produced.
Base model license
Kimi K3 License, as published by the model author on Hugging Face. The full license text ships inside the kit.
This package contains zero weight bytes. It ships the serving stack, not the model. The kit pins moonshotai/Kimi-K3 on Hugging Face at an exact commit, and on first deploy your node pulls the weights directly from https://huggingface.co/moonshotai/Kimi-K3 at that pinned commit, roughly 1.56 TB across 96 shards, under the Kimi K3 License. The repository reported as not gated at our capture time; if Moonshot enables gating you will need a Hugging Face account with the license accepted before the pull. Checksums are verified at deploy time against that pinned revision, so you can confirm the bytes you downloaded are the bytes we measured. We do not store, mirror, re-upload or redistribute these weights, and no path in this package points at a RightNow AI copy of them. Budget the download time and the disk on your first run.
The weights carry the Kimi K3 License (Copyright (c) 2026 Moonshot AI), which permits commercial use, modification, redistribution and fine-tuning, and imposes two threshold obligations you should read before you deploy. Section 3 requires prominent "Kimi K3" attribution above 100 million monthly active users or 20 million USD in monthly revenue. Section 2 is the one that binds earlier: if you or any affiliate operate Model-as-a-Service, meaning you give a third party access to inference or fine-tuning in a way that lets them exercise meaningful control over inputs, parameters or training data, and your aggregate revenue exceeds 20 million USD over any consecutive 12 months, you must enter a separate agreement with Moonshot AI before any commercial use. Note that the trigger is total company revenue, not the revenue of the service line. Section 4 carves out purely internal use and use via Moonshot's own products or certified inference partners. This is a contract gate, not a prohibition. License questions go to license@moonshot.ai, and re-read the license text at each K3 point release, because the terms changed between K2 and K3.
Serving uses public-dependency backends. The routed-expert MXFP4 layers run through the Marlin MXFP4 mixture-of-experts backend and KDA prefill runs through the Triton KDA chain, both part of public vLLM and its public dependencies. Moonshot AI's private CUDA extensions are not required by any path in this package, and the pinned engine wheel is built from the public vLLM source tree with the in-tree FlashKDA CUDA extension disabled.
Kimi K3 does not fit on an 8x B200 node, and we measured that rather than assuming it. At tensor parallel 8 with expert parallelism the weights alone allocate 175.40 GiB per GPU against 175.75 GiB free of 178.35 GiB total, which leaves no headroom for KV cache, activations, CUDA graphs or the vision tower. Four configurations were attempted and all four failed, including one with 20 GB of CPU offload requested and 256 GB of host RAM, where the offload flag proved silently ineffective on the mixture-of-experts path and the allocation was byte-identical. Serving this model requires more than 1.5 TB of HBM inside a single NVLink domain. This package targets an 8x B300 node (288 GB per GPU, 2304 GB per node). A GB200 NVL72 domain or a two-node 16x B200 NVLink domain also clears the floor but is not the configuration these numbers were measured on.
Kimi K3 is a multimodal model. It ships a MoonViT-V2 vision tower of roughly 401 million parameters and its Hugging Face pipeline is image-text-to-text. Every number in this package was measured on text-only workloads, the vision tower was not exercised, and the accuracy reasoning covers text generation only. Nothing here characterizes image input latency, throughput, memory or quality. If your traffic includes images, treat this package as uncharacterized for that path and measure it yourself.
The 2.8 trillion parameter figure is the total. K3 is a mixture-of-experts model that activates about 104 billion parameters per token, 16 routed experts of 896 plus 2 shared experts, across 92 mixture-of-experts layers. The total governs how much memory you must provision; the active count governs how fast it runs. Both matter and they answer different questions.
Every published number was measured on a single 8x B300 NVLink node at the pinned configuration (tensor parallel 8 with expert parallelism, 40960 context). The committed request-level evidence covers workload A coding prompts of about 34 to 36K tokens as reported by the server, with generation capped at 256 output tokens per request, at concurrency 1 and 8, 3 measured requests per point after 1 warmup (results/raw/serving-c1-c8-runA.json and serving-c1-c8-runB.json, two independent same-configuration runs whose throughput agrees within 0.4 percent). Throughput fields are stated per GPU; the raw files also record the node aggregate, which is 8 times the per-GPU figure. A wider sweep at concurrency 32 and 128 with 2048-token outputs completed with every request ok and bench exit code 0, but its raw per-request file was not persisted, so none of its figures are published here. The model supports a 1 million token context window; we did not measure anywhere near it. Longer context, longer outputs, other GPUs, multi-node topologies and other concurrencies are not characterized.
Peak VRAM is not published. The only figure a run records is the serving preallocation at whatever memory utilization the engine ran at, which says nothing about the serving stack itself. Size your hardware from the measured weight residency instead: 175.40 GiB per GPU at tensor parallel 8 with expert parallelism, before any KV cache.
The engine build these numbers were measured on is a source-built wheel, vLLM 0.23.1rc1.dev1475+gf68f4fdde, compiled from vllm-project/vllm at commit f68f4fdd (the head of pull request 50000) with the in-tree FlashKDA CUDA extension disabled, built and run on Modal. The kit does not ship that wheel binary. It ships the pinned build recipe for it (deploy/build/Dockerfile.vllm-k3-wheel, same commit, same exclusions), transcribed from the proven Modal build; that recipe had not yet been executed on a Docker-capable machine at verification time, so treat your first build from it as a validation run.
Version 2 is planned and scoped: when vLLM pull request 50000 merges upstream and the FlashKDA CUDA path becomes publicly buildable, v2 re-pins to the merged build, re-measures the same request grid on 8x B300, and publishes the re-measured numbers, including the FlashKDA prefill path this version runs without. Buyers of v1 receive v2 through the catalog's versioned update and maintenance subscription mechanism.
Calibration
No contamination risk, structurally. This package does not quantize anything. The MXFP4 weights are Moonshot's own quantization-aware-training release and are served exactly as published, byte for byte, which also means the accuracy reference for this package is that MXFP4 build rather than a full-precision upcast we produced. No stage in the k3-turbo stack reads a calibration corpus at any point, so overlap between calibration data and evaluation data is impossible rather than merely unlikely.
Cost basis
Cost inputs current as of 2026-07-29: no cost per token is published, because no sourced 8x B300 node hourly rate was recorded beside the measurements. Compute cost from the published throughput at your own node rate, and note the denominator is the whole 8-GPU node, not one GPU.