RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Engine comparison

vLLM vs TensorRT-LLM: a measured engine comparison

We measured Llama 3.1 8B Instruct on NVIDIA H100 80GB and NVIDIA L40S 48GB at BF16 and FP8, using vLLM 0.23.0 and TensorRT-LLM 1.2.1, as of June 20, 2026. Every claim below stays inside those conditions.

Our measured latency and throughput leads change with concurrency, while backend scope limits every comparison.

Head-to-head evidence

We compare measured throughput and TTFT p50 at every published concurrency, then show derived cost cells at each published configuration. Every ratio sits beside both source values.

Measured throughput and TTFT use render-derived ratios from the two absolute values. Cost remains labeled derived from the recorded GPU rate and is compared as published absolute values. This dataset does not define a cross-engine cost ratio.
Metric and conditionvLLMTensorRT-LLMConditional verdict
Saturation throughput on NVIDIA H100 80GB at BF16NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off

5,333 output tokens per second

4,813 output tokens per second

We measured higher saturation output throughput for vLLM. vLLM recorded 5,333 output tokens per second versus TensorRT-LLM at 4,813 output tokens per second, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.11x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off

TTFT p50 at concurrency 1NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

35 milliseconds

32 milliseconds

We measured lower TTFT p50 for TensorRT-LLM. TensorRT-LLM recorded 32 milliseconds versus vLLM at 35 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.09x.derived at render time from the two measured latency values at concurrency 1, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

TTFT p50 at concurrency 8NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

177 milliseconds

56 milliseconds

We measured lower TTFT p50 for TensorRT-LLM. TensorRT-LLM recorded 56 milliseconds versus vLLM at 177 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 3.16x.derived at render time from the two measured latency values at concurrency 8, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

TTFT p50 at concurrency 32NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

514 milliseconds

235 milliseconds

We measured lower TTFT p50 for TensorRT-LLM. TensorRT-LLM recorded 235 milliseconds versus vLLM at 514 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 2.19x.derived at render time from the two measured latency values at concurrency 32, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

TTFT p50 at concurrency 64NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

790 milliseconds

504 milliseconds

We measured lower TTFT p50 for TensorRT-LLM. TensorRT-LLM recorded 504 milliseconds versus vLLM at 790 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.57x.derived at render time from the two measured latency values at concurrency 64, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

TTFT p50 at concurrency 128NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

1,658 milliseconds

1,691 milliseconds

We measured lower TTFT p50 for vLLM. vLLM recorded 1,658 milliseconds versus TensorRT-LLM at 1,691 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.02x.derived at render time from the two measured latency values at concurrency 128, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

TTFT p50 at concurrency 256NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

2,221 milliseconds

2,848 milliseconds

We measured lower TTFT p50 for vLLM. vLLM recorded 2,221 milliseconds versus TensorRT-LLM at 2,848 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.28x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, BF16, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured

$0.206 BF16, derived

$0.228 BF16, derived

We derived lower cost for vLLM. vLLM $0.206 versus TensorRT-LLM $0.228, under NVIDIA H100 80GB, BF16, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026.

USD per 1M output tokens, derived from the recorded GPU rateNVIDIA H100 80GB, FP8, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured

$0.158 FP8, derived

TensorRT-LLM at fp8 was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100.

We do not calculate a ratio because one published cost cell is absent.

USD per 1M output tokens, derived from the recorded GPU rateNVIDIA L40S 48GB, BF16, source saturation concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured

$0.425 BF16, derived

TensorRT-LLM on the L40S was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100.

We do not calculate a ratio because one published cost cell is absent.

Throughput at saturation

We keep measured throughput on its own scale. Derived cost never shares this chart.

vLLM5,333 output tokens per second
TensorRT-LLM4,813 output tokens per second
output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. This figure uses a throughput-only zero-based scale; derived cost is not charted here.

Prefix caching depends on workload

We separate one shared prefix from many distinct prefixes because the source scopes those workloads differently.

No TensorRT-LLM prefix-cache result. Both prefix experiments ran vLLM and SGLang only.

The limitations bound every result

TensorRT-LLM ran the PyTorch backend (no engine compile), so there is no compile time to report and its peak number may differ from a built TensorRT engine.

  • We benchmarked TensorRT-LLM on its PyTorch backend, which has no ahead-of-time engine compile, so its peak throughput here is not the ceiling a compiled TensorRT engine would reach. And our single-prefix cache test is not the many-distinct-prefix case SGLang's RadixAttention is designed for, so we ran that separately too. We did not let either feed the headline.
  • One H100 80GB and one L40S 48GB, single GPU, tensor-parallel size 1.
  • Same Llama-3.1-8B-Instruct weights and same request stream across engines. Three timed repeats; the 95 percent confidence intervals are tight and live in the committed raw data.
  • Latency here is TTFT and per-request percentiles, not goodput at a fixed SLO. Whether 514 ms first-token at high load is acceptable depends on your use case; an SLO-goodput sweep is future work.
  • Numbers are self-reported on our harness. We do not cross-calibrate against an external suite like MLPerf, so the open harness is the check: re-run it and compare.
  • So the 4,813 tok/s we measured is the PyTorch-backend peak, not TensorRT-LLM's ceiling, and we got no compile-time number.
  • The L40S sweep stopped at concurrency 128 (its saturation), the H100 at 256.
  • Engines move weekly. These numbers are vLLM 0.23.0, SGLang 0.5.13, TensorRT-LLM 1.2.1, as of June 20 2026.
  • A single shared prefix is the easy case that any block-level cache handles well. It is not the workload SGLang's RadixAttention is built for.
  • RadixAttention may pull ahead with longer prefixes, deeper trees, or heavier eviction pressure than we tested. On our test it did not, and we are not going to claim otherwise.

What is not measured

We did not measure models beyond Llama 3.1 8B Instruct, GPUs beyond NVIDIA H100 80GB and NVIDIA L40S 48GB, or engine versions newer than vLLM 0.23.0 and TensorRT-LLM 1.2.1 on the June 20, 2026 as-of date.

  • No fp8 throughput and no fp8 latency. The source publishes fp8 only as a derived cost per 1M output tokens at saturation, so no fp8 tokens-per-second, first-token latency, or speedup may be rendered from this dataset.
  • No L40S throughput and no L40S latency. The L40S appears only as a derived cost cell. Every sweep row here is the H100 at bf16.
  • No TensorRT-LLM prefix-cache result. Both prefix experiments ran vLLM and SGLang only.
  • No compiled TensorRT engine. Every TensorRT-LLM number here is its PyTorch backend, which the source states is not TensorRT-LLM's throughput ceiling, and no engine compile time was recorded.
  • No prefix-cache latency values. The source states one engine had lower first-token latency on the many-prefix workload but publishes no latency numbers for either prefix experiment, so that comparison stays qualitative.
  • No hardware or precision label for the single-shared-prefix sweep. Its caption names only the concurrency, so those rows must not be presented under a GPU or a precision the source does not state.
  • No goodput at a fixed service-level objective, and no accuracy or output-quality comparison between engines. This dataset is throughput, first-token latency, and derived cost only.
  • No cross-engine or cross-precision ratio is stored. Pages derive every multiplier and percentage from the absolute pair at render time and print both absolute values beside it.
  • Not cross-calibrated against an external benchmark suite. These are self-reported numbers from an open harness, and the harness is the check.
  • No claim that these results still hold. The dataset is pinned to one as-of date and one engine version triple, and engines change on a weekly cadence.
  • TensorRT-LLM at fp8 was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100.
  • TensorRT-LLM on the L40S was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100.

Questions answered from the measurement

Which engine had higher measured saturation throughput, vLLM or TensorRT-LLM?
We measured higher saturation throughput for vLLM. vLLM recorded 5,333 output tokens per second; TensorRT-LLM recorded 4,813 output tokens per second. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026.
Which engine had lower measured TTFT p50 at the NVIDIA H100 80GB saturation point?
We measured lower TTFT p50 for vLLM. vLLM recorded 2,221 milliseconds; TensorRT-LLM recorded 2,848 milliseconds. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026.
Which engine had lower derived cost per output token in the published NVIDIA H100 80GB BF16 configuration?
We derived lower cost for vLLM. vLLM was $0.206; TensorRT-LLM was $0.228 USD per 1M output tokens. Condition: NVIDIA H100 80GB, BF16, source saturation concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. Both values are derived from the recorded GPU rate and measured saturation throughput.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy