qwen3-6-27b-fp8cd-v3-mlponly-h100-vllm
Published Jul 25, 2026. Proof verified Jul 25, 2026.
1.28x faster
Median latency measured 2857 ms at baseline and 2215 ms optimized.
Quality evidence
gsm8k 99.87% recovery, passed.
RunInfra has published 1 measured package for Qwen/Qwen3.6-27B on H100. Every package below keeps its engine, price, and proof date together.
| Package | GPU | Engine | Optimization | Price | Verified |
|---|---|---|---|---|---|
| qwen3-6-27b-fp8cd-v3-mlponly-h100-vllm | H100 | vLLM 0.25.1 | Channelwise FP8, measured selective-layer recipe | $40 one-time | Jul 25, 2026 |
Selective channelwise FP8 leaves the sensitive paths at higher precision while converting the remaining eligible layers.
Published Jul 25, 2026. Proof verified Jul 25, 2026.
Median latency measured 2857 ms at baseline and 2215 ms optimized.
gsm8k 99.87% recovery, passed.
Published package measurements and prices for Qwen3.6 27B, grouped from the public RunInfra catalog.
Measured by RunInfra on H100 with vLLM 0.25.1; proof dates are listed with each package.
We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.
Other GPU measurements for this model not published.
Other engine measurements for this model not published.
Measurements for context lengths not listed in these packages not published.
qwen3-6-27b-fp8cd-v3-mlponly-h100-vllm is listed at $40 one-time, published Jul 25, 2026.
qwen3-6-27b-fp8cd-v3-mlponly-h100-vllm is measured on H100 with vLLM 0.25.1. Proof verified Jul 25, 2026.
gsm8k 99.87% recovery, passed. Proof verified Jul 25, 2026.
© 2026 RunInfra. All rights reserved.