fp8_dynamic quantized and serving-tuned for A100-40GB on vllm v0.18.0, with the benchmark receipt included. You get a self-contained kit: the optimized weights, the exact serving configuration, and the measured proof. Run it on any cloud or your own hardware. Nothing calls home, nothing expires.
Measured against the unoptimized baseline on A100-40GB: median latency 327 ms down to 88 ms. The full benchmark methodology and receipt ship inside the kit.
The weights run on your infrastructure. No deprecation notices, no silent model swaps, no provider deciding what your product is allowed to do.
One purchase instead of a meter that runs forever. A closed API charges you for every token, at whatever the price becomes; your own deployment costs what your GPU costs.
Prompts and outputs never leave your network. Nothing calls home, nothing is retained by someone else, and compliance stops being a negotiation.
vllm v0.18.0, fp8_dynamic, tuned for A100-40GB (sm_80). Requires 40 GB VRAM minimum.
Apache-2.0, as published by the model author on Hugging Face. The full license text ships inside the kit.
The kit is self-contained: bring it to any cloud with A100-40GB-class GPUs, or your own hardware.
© 2026 RunInfra. All rights reserved.
Describe the goal. RunInfra builds and optimizes the stack.