Engine comparison
We measured Llama 3.1 8B Instruct on NVIDIA H100 80GB and NVIDIA L40S 48GB at BF16 and FP8, using vLLM 0.23.0 and SGLang 0.5.13, as of June 20, 2026. Every claim below stays inside those conditions.
Our measurements split the lead by condition, making workload shape more useful than a universal ranking.
We compare measured throughput and TTFT p50 at every published concurrency, then show derived cost cells at each published configuration. Every ratio sits beside both source values.
| Metric and condition | vLLM | SGLang | Conditional verdict |
|---|---|---|---|
| Saturation throughput on NVIDIA H100 80GB at BF16NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off | 5,333 output tokens per second | 5,235 output tokens per second | We measured higher saturation output throughput for vLLM. vLLM recorded 5,333 output tokens per second versus SGLang at 5,235 output tokens per second, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.02x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off |
| TTFT p50 at concurrency 1NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 35 milliseconds | 41 milliseconds | We measured lower TTFT p50 for vLLM. vLLM recorded 35 milliseconds versus SGLang at 41 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.17x.derived at render time from the two measured latency values at concurrency 1, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 8NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 177 milliseconds | 221 milliseconds | We measured lower TTFT p50 for vLLM. vLLM recorded 177 milliseconds versus SGLang at 221 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.25x.derived at render time from the two measured latency values at concurrency 8, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 32NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 514 milliseconds | 597 milliseconds | We measured lower TTFT p50 for vLLM. vLLM recorded 514 milliseconds versus SGLang at 597 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.16x.derived at render time from the two measured latency values at concurrency 32, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 64NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 790 milliseconds | 997 milliseconds | We measured lower TTFT p50 for vLLM. vLLM recorded 790 milliseconds versus SGLang at 997 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.26x.derived at render time from the two measured latency values at concurrency 64, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 128NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 1,658 milliseconds | 1,761 milliseconds | We measured lower TTFT p50 for vLLM. vLLM recorded 1,658 milliseconds versus SGLang at 1,761 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.06x.derived at render time from the two measured latency values at concurrency 128, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 256NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 2,221 milliseconds | 2,543 milliseconds | We measured lower TTFT p50 for vLLM. vLLM recorded 2,221 milliseconds versus SGLang at 2,543 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.14x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, BF16, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured | $0.206 BF16, derived | $0.21 BF16, derived | We derived lower cost for vLLM. vLLM $0.206 versus SGLang $0.21, under NVIDIA H100 80GB, BF16, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, FP8, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured | $0.158 FP8, derived | $0.17 FP8, derived | We derived lower cost for vLLM. vLLM $0.158 versus SGLang $0.17, under NVIDIA H100 80GB, FP8, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA L40S 48GB, BF16, source saturation concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured | $0.425 BF16, derived | $0.438 BF16, derived | We derived lower cost for vLLM. vLLM $0.425 versus SGLang $0.438, under NVIDIA L40S 48GB, BF16, source saturation concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026. |
We keep measured throughput on its own scale. Derived cost never shares this chart.
We separate one shared prefix from many distinct prefixes because the source scopes those workloads differently.
A single shared prefix is the easy case that any block-level cache handles well. It is not the workload SGLang's RadixAttention is built for.
We measured higher cache-on throughput for vLLM. vLLM, cache on, 90 percent hit rate recorded 2,855 output tokens per second versus vLLM, cache on, zero hit rate at 999 output tokens per second, under One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. As of June 20, 2026.. The render-derived ratio is 2.86x.derived at render time for vLLM from its own cache-on throughput at a 90 percent hit rate against its cache-on throughput at a zero hit rate, output tokens per second, mean of three timed repeats after warmup, cache state per row
We measured higher cache-on throughput for SGLang. SGLang, cache on, 90 percent hit rate recorded 2,488 output tokens per second versus SGLang, cache on, zero hit rate at 937 output tokens per second, under One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. As of June 20, 2026.. The render-derived ratio is 2.66x.derived at render time for SGLang from its own cache-on throughput at a 90 percent hit rate against its cache-on throughput at a zero hit rate, output tokens per second, mean of three timed repeats after warmup, cache state per row
RadixAttention may pull ahead with longer prefixes, deeper trees, or heavier eviction pressure than we tested. On our test it did not, and we are not going to claim otherwise.
We measured higher measured throughput on the many-prefix workload. vLLM recorded 2,677 output tokens per second versus SGLang at 2,493 output tokens per second, under 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. Distinct prefixes: 8. NVIDIA H100 80GB. Precision was not published. As of June 20, 2026.. The render-derived ratio is 1.07x.derived at render time at 8 distinct prefixes, output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on
We measured higher measured throughput on the many-prefix workload. vLLM recorded 2,277 output tokens per second versus SGLang at 2,086 output tokens per second, under 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. Distinct prefixes: 32. NVIDIA H100 80GB. Precision was not published. As of June 20, 2026.. The render-derived ratio is 1.09x.derived at render time at 32 distinct prefixes, output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on
We measured higher measured throughput on the many-prefix workload. vLLM recorded 1,467 output tokens per second versus SGLang at 1,376 output tokens per second, under 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. Distinct prefixes: 128. NVIDIA H100 80GB. Precision was not published. As of June 20, 2026.. The render-derived ratio is 1.07x.derived at render time at 128 distinct prefixes, output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on
We did not measure models beyond Llama 3.1 8B Instruct, GPUs beyond NVIDIA H100 80GB and NVIDIA L40S 48GB, or engine versions newer than vLLM 0.23.0 and SGLang 0.5.13 on the June 20, 2026 as-of date.
© 2026 RunInfra. All rights reserved.