Serving engine
We measured throughput and TTFT for Llama 3.1 8B Instruct with TensorRT-LLM 1.2.1 on the NVIDIA H100 80GB at BF16, as of June 20, 2026. Derived cost configurations reference NVIDIA H100 80GB and NVIDIA L40S 48GB. Source absences render below.
0 published packages serve on this engine. Package rows keep their own engine versions and verified dates.
TensorRT-LLM's PyTorch-backend run makes backend scope part of every throughput and latency reading.
Source article and methodMeasured throughput and TTFT remain separate from derived cost. Every comparison ratio shows both source values and its full condition.
| Concurrency | Measured throughput | Measured TTFT p50 | Full condition |
|---|---|---|---|
| 1 | 152 output tokens per second | 32 milliseconds | NVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026 |
| 8 | 1,027 output tokens per second | 56 milliseconds | NVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026 |
| 32 | 2,978 output tokens per second | 235 milliseconds | NVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026 |
| 64 | 4,029 output tokens per second | 504 milliseconds | NVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026 |
| 128 | 4,774 output tokens per second | 1,691 milliseconds | NVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026 |
| 256 | 4,813 output tokens per second | 2,848 milliseconds | NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026 |
We measured higher output throughput for vLLM. vLLM recorded 5,333 output tokens per second versus TensorRT-LLM at 4,813 output tokens per second, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.11x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off
We measured lower TTFT p50 for vLLM. vLLM recorded 2,221 milliseconds versus TensorRT-LLM at 2,848 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.28x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column
We measured higher output throughput for SGLang. SGLang recorded 5,235 output tokens per second versus TensorRT-LLM at 4,813 output tokens per second, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.09x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off
We measured lower TTFT p50 for SGLang. SGLang recorded 2,543 milliseconds versus TensorRT-LLM at 2,848 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.12x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column
| Configuration | Published value or absence | Full condition |
|---|---|---|
| NVIDIA H100 80GB BF16 | $0.228 USD per 1M output tokens, derived | NVIDIA H100 80GB, BF16, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026 |
| NVIDIA H100 80GB FP8 | TensorRT-LLM at fp8 was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | NVIDIA H100 80GB, FP8, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026 |
| NVIDIA L40S 48GB BF16 | TensorRT-LLM on the L40S was not measured. The cost block publishes a single TensorRT-LLM configuration, bf16 on the H100. | NVIDIA L40S 48GB, BF16, source-selected cost concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026 |
No TensorRT-LLM prefix-cache result. Both prefix experiments ran vLLM and SGLang only.
One client drives all three engines with identical request streams, so the timing definitions are the same everywhere. Time to first token is the time to the first streamed token. Throughput is output tokens per second. Warm up, then three timed repeats per operating point, charted as the mean; the 95 percent confidence intervals are tight and live in the committed raw data. Prefix caching is off for the unique-prompt sweeps. Concurrency swept from 1 to 256 at 1,024 input and 256 output tokens, unique prompts, prefix caching off. Same weights and same request stream across engines, single GPU, tensor-parallel size 1. 0 published package rows retain their recorded engine version, concurrency, and verified date.
We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.
These rows come from the published catalog, not the comparison sweep. Each package retains its own model, engine version, concurrency, and verified date.
No published package currently serves on this engine.
© 2026 RunInfra. All rights reserved.