Optimization technique
Selective-layer channelwise FP8 keeps the published sensitive paths at higher precision while quantizing the remaining eligible layers.
We measured 1 published package across 1 GPU target, using the recorded baseline and optimized conditions.
The selective recipe exists because a measured per-layer search found which paths lose accuracy under FP8, so those stay in higher precision while the rest converts.
We show absolute baseline and optimized values before the shared verdict. We keep each condition attached to its result.
| Model | Target | Throughput | Median latency | Verdict and conditions | Accuracy | Price and date |
|---|---|---|---|---|---|---|
| Qwen3.6 27BQwen/Qwen3.6-27B | H100vLLM 0.25.1 | 361 tokens/s baseline465 tokens/s optimized | 2857 ms p50 baseline2215 ms p50 optimized | 1.28x faster.baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. | Baseline 0.5754, optimized 0.5747.gsm8k 99.87% recovery, passedStandard errors not published. | $40Verified Jul 25, 2026 |
We restate the published throughput pairs. The neutral fills do not declare a winner.
1.28x faster.
We do not call a score change an improvement when the gap is smaller than about twice the combined standard errors.
Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. The optimized score measured below the baseline, at 99.87% of it. Neither score published a standard error, so this difference cannot be separated from sampling noise. The difference is not stated as a regression. Standard errors not published.
Tasks beyond gsm8k exact_match strict (N=1319, completion protocol) not published.
GPUs beyond H100 not published.
Serving engines beyond vLLM 0.25.1 not published.
We tie each accuracy verdict to its published task metric. Qwen3.6 27B: Throughput measured 361 tokens/s at baseline and 465 tokens/s optimized. P50 latency measured 2857 ms at baseline and 2215 ms optimized. 1.28x faster. Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026.
We compare throughput and p50 latency under the recorded conditions. Qwen3.6 27B: Throughput measured 361 tokens/s at baseline and 465 tokens/s optimized. P50 latency measured 2857 ms at baseline and 2215 ms optimized. 1.28x faster. Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026.
Qwen3.6 27B are the published models on this page. Qwen3.6 27B: Throughput measured 361 tokens/s at baseline and 465 tokens/s optimized. P50 latency measured 2857 ms at baseline and 2215 ms optimized. 1.28x faster. Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026.
Qwen3.6 27B: Throughput measured 361 tokens/s at baseline and 465 tokens/s optimized. P50 latency measured 2857 ms at baseline and 2215 ms optimized. 1.28x faster. Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026. Tasks beyond gsm8k exact_match strict (N=1319, completion protocol) not published. GPUs beyond H100 not published. Serving engines beyond vLLM 0.25.1 not published.
© 2026 RunInfra. All rights reserved.