Open models, quantized and benchmarked on real GPUs. Buy once, run anywhere.
Quantized and serving-tuned for H100 on vLLM 0.25.1, with the benchmark receipt included. You get a self-contained kit: the optimized weights, the exact serving configuration, and the measured proof. Run it on any cloud or your own hardware. Nothing calls home, nothing expires.
We keep re-optimizing and re-benchmarking as engines move, and you always download the latest verified version. Cancel anytime. Nothing is charged until you confirm at checkout.
Company
© 2026 RunInfra. All rights reserved.
The kit is self-contained: bring it to any cloud with H100-class GPUs, or your own hardware.
The weights run on your infrastructure. No deprecation notices, no silent model swaps, no provider deciding what your product is allowed to do.
One purchase instead of a meter that runs forever. A closed API charges you for every token, at whatever the price becomes; your own deployment costs what your GPU costs.
Prompts and outputs never leave your network. Nothing calls home, nothing is retained by someone else, and compliance stops being a negotiation.
llm-compressor 0.12.0, data-free: no calibration corpus is read at any point, so the quantization cannot have seen evaluation data
the exact engine build and arguments every receipt number was produced on: vLLM 0.25.1, max_num_seqs 256, max_num_batched_tokens 8192, max-model-len 4096, chosen as the best of a five-point grid measured on this model rather than inherited from another package
vLLM 0.25.1, Channelwise FP8 weights, dynamic per-token activations, tuned for H100.
Measured accuracy
(N=100, chat-template reasoning-aware harness, extraction v2)
Passed accuracy gate
No measurable change in accuracy
The difference between the baseline and optimized scores is smaller than the combined measurement error, so it is within measurement noise.
Methodology
Buyer disclosure
Two licenses apply: the RunInfra package license you purchase under, and the base model's own open-source license.
RunInfra Package License v1.1 (2026-07-25)
Buy once. No meter, no expiry, nothing calling home. The full terms ship inside the kit.
This package is licensed to the purchasing workspace, one time, for the exact version purchased. Everyone in that workspace may run it in production without limits: unlimited inference, on any hardware or cloud the workspace controls, for any lawful commercial purpose. You may modify the configuration and tooling for your own use.
Everything the model produces for you belongs to you. You may use, sell, and build products on the model's outputs without restriction or royalty. Serving the model to your own customers as part of your product is use, not redistribution, and is fully allowed.
You may not redistribute, resell, sublicense, rent, publish, or otherwise make the package or its artifacts available to any third party. That covers the optimized weights, serving configuration, scripts, kit archive, and benchmark receipts, whole or in part, modified or not. The license belongs to the purchasing workspace and cannot be transferred separately from it.
The underlying model remains under its upstream open-source license, which is included in this kit with attribution and a statement of RunInfra's modifications. Nothing in this license restricts rights the upstream license grants you for the ORIGINAL model; the restrictions above apply to RunInfra's optimized package.
You purchased this exact version, proven against the exact engine version named in the receipt. It stays downloadable to your workspace and never expires. The benchmark receipt describes measurements taken at verification time on the named hardware; RunInfra does not promise future updates to this version, and later package versions are separate purchases.
If the workspace redistributes the package or its artifacts, this license terminates for that workspace. Sections about your outputs survive termination for outputs already produced.
Apache-2.0, as published by the model author on Hugging Face. The full license text ships inside the kit.