Open models, quantized and benchmarked on real GPUs. Buy once, run anywhere.
Channelwise FP8 weights, dynamic per-token activations quantized and serving-tuned for H100 on vLLM 0.25.1, with the benchmark receipt included. You get a self-contained kit: the optimized weights, the exact serving configuration, and the measured proof. Run it on any cloud or your own hardware. Nothing calls home, nothing expires.
We keep re-optimizing and re-benchmarking as engines move, and you always download the latest verified version. Cancel anytime. Nothing is charged until you confirm at checkout.
© 2026 RunInfra. All rights reserved.
Measured against the unoptimized baseline on H100: median latency 1058 ms down to 815 ms. Concurrency 8. The full benchmark methodology and receipt ship inside the kit.
The kit is self-contained: bring it to any cloud with H100-class GPUs, or your own hardware.
The weights run on your infrastructure. No deprecation notices, no silent model swaps, no provider deciding what your product is allowed to do.
One purchase instead of a meter that runs forever. A closed API charges you for every token, at whatever the price becomes; your own deployment costs what your GPU costs.
Prompts and outputs never leave your network. Nothing calls home, nothing is retained by someone else, and compliance stops being a negotiation.
llm-compressor 0.12.0, data-free: no calibration corpus is read at any point
the ignore list is re:.*visual.* and re:.*linear_attn.*, carried over from the measured per-layer sensitivity map for this qwen3_5 hybrid family; this is the third model built on that map, and it was validated here by the full-set gsm8k run rather than re-searched layer by layer
the exact engine build and arguments the receipt numbers were produced on: vLLM 0.25.1, max-model-len 4096, max-num-seqs 64, gpu-memory-utilization 0.90, speculative decoding off
vLLM 0.25.1, Channelwise FP8 weights, dynamic per-token activations, tuned for H100.
Measured accuracy
(N=1319, completion protocol)
Passed accuracy gate
Measured difference, significance not published
The optimized score measured below the baseline, at 99.35% of it. Neither score published a standard error, so this difference cannot be separated from sampling noise. The difference is not stated as a regression.
Methodology
Buyer disclosure
Two licenses apply: the RunInfra package license you purchase under, and the base model's own open-source license.
Buy once and the kit is yours forever: run it on any cloud or your own hardware. No meter, no expiry, nothing calling home. The full terms ship inside the kit.
Apache-2.0, as published by the model author on Hugging Face. The full license text ships inside the kit.