Optimization technique
Under H100, vLLM 0.25.1, request profile chat-128, concurrency 8, measured basis, verified Jul 25, 2026, Qwen3.6 27B with Channelwise FP8, measured selective-layer recipe: 1.28x faster; 2857 ms at baseline and 2215 ms optimized. This page covers 1 published package using this technique and keeps every claim tied to its recorded run. Other tasks, GPUs, and serving engines are not published.
Selective-layer channelwise FP8 keeps the published sensitive paths at higher precision while quantizing the remaining eligible layers.
We measured 1 published package across 1 GPU target, using the recorded baseline and optimized conditions.
The selective recipe exists because a measured per-layer search found which paths lose accuracy under FP8, so those stay in higher precision while the rest converts.
Measurement policy: Methodology.
We show absolute baseline and optimized values before the shared verdict. We keep each condition attached to its result.
| Model | Target | Throughput | Median latency | Verdict and conditions | Accuracy | Price and date |
|---|---|---|---|---|---|---|
| Qwen3.6 27BQwen/Qwen3.6-27B | H100vLLM 0.25.1 | 361 tokens/s baseline465 tokens/s optimized | 2857 ms p50 baseline2215 ms p50 optimized | 1.28x faster.baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. | Baseline 0.5754, optimized 0.5747.gsm8k 99.87% recovery, passedStandard errors not published. | $40Verified Jul 25, 2026 |
We restate the published throughput pairs. The neutral fills do not declare a winner.
1.28x faster.
We do not call a score change an improvement when the gap is smaller than about twice the combined standard errors.
Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. The optimized score measured below the baseline, at 99.87% of it. Neither score published a standard error, so this difference cannot be separated from sampling noise. The difference is not stated as a regression. Standard errors not published.
Tasks beyond gsm8k exact_match strict (N=1319, completion protocol) not published.
GPUs beyond H100 not published.
Serving engines beyond vLLM 0.25.1 not published.
We tie each accuracy verdict to its published task metric. Qwen3.6 27B: Throughput measured 361 tokens/s at baseline and 465 tokens/s optimized. P50 latency measured 2857 ms at baseline and 2215 ms optimized. 1.28x faster. Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026.
We compare throughput and p50 latency under the recorded conditions. Qwen3.6 27B: Throughput measured 361 tokens/s at baseline and 465 tokens/s optimized. P50 latency measured 2857 ms at baseline and 2215 ms optimized. 1.28x faster. Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026.
Qwen3.6 27B are the published models on this page. Qwen3.6 27B: Throughput measured 361 tokens/s at baseline and 465 tokens/s optimized. P50 latency measured 2857 ms at baseline and 2215 ms optimized. 1.28x faster. Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026.
Qwen3.6 27B: Throughput measured 361 tokens/s at baseline and 465 tokens/s optimized. P50 latency measured 2857 ms at baseline and 2215 ms optimized. 1.28x faster. Qwen3.6 27B measured 0.5754 at baseline and 0.5747 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026. Tasks beyond gsm8k exact_match strict (N=1319, completion protocol) not published. GPUs beyond H100 not published. Serving engines beyond vLLM 0.25.1 not published.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs