fp8_channel_dynamic_s2 quantized and serving-tuned for H100 on vLLM 0.25.1, with the benchmark receipt included. You get a self-contained kit: the optimized weights, the exact serving configuration, and the measured proof. Run it on any cloud or your own hardware. Nothing calls home, nothing expires.
Measured against the unoptimized baseline on H100: median latency 2857 ms down to 2215 ms. The full benchmark methodology and receipt ship inside the kit.
The weights run on your infrastructure. No deprecation notices, no silent model swaps, no provider deciding what your product is allowed to do.
One purchase instead of a meter that runs forever. A closed API charges you for every token, at whatever the price becomes; your own deployment costs what your GPU costs.
Prompts and outputs never leave your network. Nothing calls home, nothing is retained by someone else, and compliance stops being a negotiation.
vLLM 0.25.1, fp8_channel_dynamic_s2, tuned for H100 (sm_90). Requires 80 GB VRAM minimum.
Apache-2.0, as published by the model author on Hugging Face. The full license text ships inside the kit.
The kit is self-contained: bring it to any cloud with H100-class GPUs, or your own hardware.
© 2026 RunInfra. All rights reserved.
Describe the goal. RunInfra builds and optimizes the stack.