qwythos-9b-fp8cd-gdnvis-h100-vllm
Published Jul 25, 2026. Proof verified Jul 25, 2026.
1.29x faster
Median latency measured 1058 ms at baseline and 815 ms optimized.
Quality evidence
gsm8k 99.35% recovery, passed.
RunInfra lists 1 measured package for Qwythos-9B-Claude-Mythos-5-1M. The newest published row, qwythos-9b-fp8cd-gdnvis-h100-vllm, costs $20 one-time on H100 with vLLM 0.25.1. Its speed result is 1.29x faster. Request profile: chat-128. Concurrency 8. That result applies only to those recorded conditions. Proof verified Jul 25, 2026. Other package prices and measurements remain separate below.
| Package | GPU | Engine | Optimization | Price | Verified |
|---|---|---|---|---|---|
| qwythos-9b-fp8cd-gdnvis-h100-vllm | H100 | vLLM 0.25.1 | Channelwise FP8 weights, dynamic per-token activations | $20 one-time | Jul 25, 2026 |
The recipe keeps the model's sensitive recurrent and vision paths at higher precision while applying channelwise FP8 elsewhere.
Published Jul 25, 2026. Proof verified Jul 25, 2026.
Median latency measured 1058 ms at baseline and 815 ms optimized.
gsm8k 99.35% recovery, passed.
Published package measurements and prices for Qwythos-9B-Claude-Mythos-5-1M, grouped from the public RunInfra catalog.
Measured by RunInfra on H100 with vLLM 0.25.1; proof dates are listed with each package.
We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/benchmarks.
Other GPU measurements for this model not published.
Other engine measurements for this model not published.
Measurements for context lengths not listed in these packages not published.
qwythos-9b-fp8cd-gdnvis-h100-vllm is listed at $20 one-time, published Jul 25, 2026.
qwythos-9b-fp8cd-gdnvis-h100-vllm is measured on H100 with vLLM 0.25.1. Proof verified Jul 25, 2026.
gsm8k 99.35% recovery, passed. Proof verified Jul 25, 2026.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs