Engine comparison
We measured Llama 3.1 8B Instruct on NVIDIA H100 80GB and NVIDIA L40S 48GB at BF16 and FP8, using vLLM 0.23.0 and TensorRT-LLM 1.2.1, as of June 20, 2026. Every claim below stays inside those conditions.
Our measured latency and throughput leads change with concurrency, while backend scope limits every comparison.
We compare measured throughput and TTFT p50 at every published concurrency, then show derived cost cells at each published configuration. Every ratio sits beside both source values.
| Metric and condition | vLLM | TensorRT-LLM | Conditional verdict |
|---|---|---|---|
| Saturation throughput on NVIDIA H100 80GB at BF16NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off | 5,333 output tokens per second | 4,813 output tokens per second | We measured higher saturation output throughput for vLLM. vLLM recorded 5,333 output tokens per second versus TensorRT-LLM at 4,813 output tokens per second, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.11x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off |
| TTFT p50 at concurrency 1NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 35 milliseconds | 32 milliseconds | We measured lower TTFT p50 for TensorRT-LLM. TensorRT-LLM recorded 32 milliseconds versus vLLM at 35 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.09x.derived at render time from the two measured latency values at concurrency 1, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 8NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 177 milliseconds | 56 milliseconds | We measured lower TTFT p50 for TensorRT-LLM. TensorRT-LLM recorded 56 milliseconds versus vLLM at 177 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 3.16x.derived at render time from the two measured latency values at concurrency 8, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 32NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 514 milliseconds | 235 milliseconds | We measured lower TTFT p50 for TensorRT-LLM. TensorRT-LLM recorded 235 milliseconds versus vLLM at 514 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 2.19x.derived at render time from the two measured latency values at concurrency 32, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 64NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 790 milliseconds | 504 milliseconds | We measured lower TTFT p50 for TensorRT-LLM. TensorRT-LLM recorded 504 milliseconds versus vLLM at 790 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.57x.derived at render time from the two measured latency values at concurrency 64, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 128NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 1,658 milliseconds | 1,691 milliseconds | We measured lower TTFT p50 for vLLM. vLLM recorded 1,658 milliseconds versus TensorRT-LLM at 1,691 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.02x.derived at render time from the two measured latency values at concurrency 128, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 256NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 2,221 milliseconds | 2,848 milliseconds | We measured lower TTFT p50 for vLLM. vLLM recorded 2,221 milliseconds versus TensorRT-LLM at 2,848 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.28x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, BF16, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured | $0.206 BF16, derived | $0.228 BF16, derived | We derived lower cost for vLLM. vLLM $0.206 versus TensorRT-LLM $0.228, under NVIDIA H100 80GB, BF16, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, FP8, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured | $0.158 FP8, derived | TensorRT-LLM at fp8 was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | We do not calculate a ratio because one published cost cell is absent. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA L40S 48GB, BF16, source saturation concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured | $0.425 BF16, derived | TensorRT-LLM on the L40S was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | We do not calculate a ratio because one published cost cell is absent. |
We keep measured throughput on its own scale. Derived cost never shares this chart.
We separate one shared prefix from many distinct prefixes because the source scopes those workloads differently.
No TensorRT-LLM prefix-cache result. Both prefix experiments ran vLLM and SGLang only.
TensorRT-LLM ran the PyTorch backend (no engine compile), so there is no compile time to report and its peak number may differ from a built TensorRT engine.
We did not measure models beyond Llama 3.1 8B Instruct, GPUs beyond NVIDIA H100 80GB and NVIDIA L40S 48GB, or engine versions newer than vLLM 0.23.0 and TensorRT-LLM 1.2.1 on the June 20, 2026 as-of date.
© 2026 RunInfra. All rights reserved.