Engine comparison
H100 BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026: vLLM 5,333 versus TensorRT-LLM 4,813 output tokens per second. TensorRT-LLM recorded lower TTFT at four of six concurrencies (1, 8, 32, and 64), vLLM at 128 and 256. Same condition: vLLM 2,221 versus TensorRT-LLM 2,848 milliseconds TTFT p50. Derived cost, same condition: vLLM $0.206 versus TensorRT-LLM $0.228 per million output tokens. From three repeats, confidence intervals unpublished. No universal winner is claimed.
Throughput and TTFT were measured on NVIDIA H100 80GB at BF16. Cost rows are derived for the published H100 BF16, H100 FP8, and L40S BF16 configurations. The comparison uses Llama 3.1 8B Instruct, vLLM 0.23.0 and TensorRT-LLM 1.2.1, as of June 20, 2026. RunInfra's later pages record vLLM 0.25.1 in published model packages (read September 21, 2026). Those newer versions were not compared in this June sweep. Sources: published model packages.
Our measured latency and throughput leads change with concurrency, while backend scope limits every comparison.
We compare measured throughput and TTFT p50 at every published concurrency, then show derived cost cells at each published configuration. Every ratio sits beside both source values.
| Metric and condition | vLLM | TensorRT-LLM | Conditional verdict |
|---|---|---|---|
| Highest measured throughput on NVIDIA H100 80GB at BF16NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off | 5,333 output tokens per second | 4,813 output tokens per second | vLLM recorded 5,333 output tokens per second versus TensorRT-LLM at 4,813 output tokens per second, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The recorded difference is 520 output tokens per second (9.8% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The render-derived ratio is 1.11x, an arithmetic comparison only.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off |
| TTFT p50 at concurrency 1NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 35 milliseconds | 32 milliseconds | vLLM recorded 35 milliseconds versus TensorRT-LLM at 32 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026. The recorded difference is 3 milliseconds (8.6% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. |
| TTFT p50 at concurrency 8NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 177 milliseconds | 56 milliseconds | vLLM recorded 177 milliseconds versus TensorRT-LLM at 56 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026. The recorded difference is 121 milliseconds (68.4% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The render-derived ratio is 3.16x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 8, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 32NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 514 milliseconds | 235 milliseconds | vLLM recorded 514 milliseconds versus TensorRT-LLM at 235 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026. The recorded difference is 279 milliseconds (54.3% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The render-derived ratio is 2.19x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 32, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 64NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 790 milliseconds | 504 milliseconds | vLLM recorded 790 milliseconds versus TensorRT-LLM at 504 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026. The recorded difference is 286 milliseconds (36.2% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The render-derived ratio is 1.57x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 64, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| TTFT p50 at concurrency 128NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 1,658 milliseconds | 1,691 milliseconds | vLLM recorded 1,658 milliseconds versus TensorRT-LLM at 1,691 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026. The recorded difference is 33 milliseconds (2.0% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. |
| TTFT p50 at concurrency 256NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column | 2,221 milliseconds | 2,848 milliseconds | vLLM recorded 2,221 milliseconds versus TensorRT-LLM at 2,848 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The recorded difference is 627 milliseconds (22.0% of the larger value), from three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. The render-derived ratio is 1.28x, an arithmetic comparison only.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, BF16, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026Derived from the recorded GPU rate and measured throughput at the source-selected cost concurrency, not measured directly. | $0.206 BF16, derived | $0.228 BF16, derived | Published derived costs: vLLM $0.206 versus TensorRT-LLM $0.228, under NVIDIA H100 80GB, BF16, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The gap between these rounded derived costs is 9.6% of the larger cost. The published throughput inputs differ by 9.8% of the larger mean. The throughput method uses three timed repeats; confidence intervals are unpublished, so statistical significance is unknown. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, FP8, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026Derived from the recorded GPU rate and measured throughput at the source-selected cost concurrency, not measured directly. | $0.158 FP8, derived | TensorRT-LLM at fp8 was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | We do not calculate a ratio because one published cost cell is absent. |
| USD per 1M output tokens, derived from the recorded GPU rateNVIDIA L40S 48GB, BF16, source-selected cost concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026Derived from the recorded GPU rate and measured throughput at the source-selected cost concurrency, not measured directly. | $0.425 BF16, derived | TensorRT-LLM on the L40S was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | We do not calculate a ratio because one published cost cell is absent. |
We keep measured throughput on its own scale. Derived cost never shares this chart.
We separate one shared prefix from many distinct prefixes because the source scopes those workloads differently.
We did not measure TensorRT-LLM's prefix cache. Both prefix experiments ran vLLM and SGLang only.
TensorRT-LLM 1.2.1 ran the PyTorch backend, its only execution backend. NVIDIA's TensorRT-LLM 1.2 release notes document removal of the TensorRT backend and engine-build CLI.
We measured throughput and TTFT only for Llama 3.1 8B Instruct on NVIDIA H100 80GB at BF16. Cost rows are limited to H100 BF16, H100 FP8, and L40S BF16. We did not test engine versions newer than vLLM 0.23.0 and TensorRT-LLM 1.2.1 on the June 20, 2026 as-of date.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs