RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Serving engine

vLLM: what RunInfra measured

We measured throughput and TTFT for Llama 3.1 8B Instruct with vLLM 0.23.0 on the NVIDIA H100 80GB at BF16, as of June 20, 2026. Derived cost configurations reference NVIDIA H100 80GB and NVIDIA L40S 48GB. Source absences render below.

4 published packages serve on this engine. Package rows keep their own engine versions and verified dates.

vLLM changes position across load and cache shape, so each decision stays tied to its measured condition.

Source article and method

The comparison evidence stays inside one measured scope

Measured throughput and TTFT remain separate from derived cost. Every comparison ratio shows both source values and its full condition.

output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off. p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column.
ConcurrencyMeasured throughputMeasured TTFT p50Full condition
1158 output tokens per second35 millisecondsNVIDIA H100 80GB, BF16, concurrency 1, Llama 3.1 8B Instruct, as of June 20, 2026
81,096 output tokens per second177 millisecondsNVIDIA H100 80GB, BF16, concurrency 8, Llama 3.1 8B Instruct, as of June 20, 2026
322,921 output tokens per second514 millisecondsNVIDIA H100 80GB, BF16, concurrency 32, Llama 3.1 8B Instruct, as of June 20, 2026
644,040 output tokens per second790 millisecondsNVIDIA H100 80GB, BF16, concurrency 64, Llama 3.1 8B Instruct, as of June 20, 2026
1284,943 output tokens per second1,658 millisecondsNVIDIA H100 80GB, BF16, concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026
2565,333 output tokens per second2,221 millisecondsNVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026

Pairwise readings at the highest measured concurrency

  • We measured higher output throughput for vLLM. vLLM recorded 5,333 output tokens per second versus SGLang at 5,235 output tokens per second, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.02x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off

  • We measured lower TTFT p50 for vLLM. vLLM recorded 2,221 milliseconds versus SGLang at 2,543 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.14x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

  • We measured higher output throughput for vLLM. vLLM recorded 5,333 output tokens per second versus TensorRT-LLM at 4,813 output tokens per second, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.11x.derived at render time from the two measured throughput values at concurrency 256, output tokens per second, mean of three timed repeats after warmup, unique prompts with prefix caching off

  • We measured lower TTFT p50 for vLLM. vLLM recorded 2,221 milliseconds versus TensorRT-LLM at 2,848 milliseconds, under NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. The render-derived ratio is 1.28x.derived at render time from the two measured latency values at concurrency 256, p50 time to the first streamed token, charted as the mean of three timed repeats after warmup, same request stream as the throughput column

Derived cost stays on its own basis

derived from the GPU hourly price and the measured saturation throughput for that configuration, not measured. Formula: cost per 1M tokens equals the GPU hourly price divided by sustained tokens per hour.
ConfigurationPublished value or absenceFull condition
NVIDIA H100 80GB BF16$0.206 USD per 1M output tokens, derivedNVIDIA H100 80GB, BF16, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026
NVIDIA H100 80GB FP8$0.158 USD per 1M output tokens, derivedNVIDIA H100 80GB, FP8, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026
NVIDIA L40S 48GB BF16$0.425 USD per 1M output tokens, derivedNVIDIA L40S 48GB, BF16, source-selected cost concurrency 128, Llama 3.1 8B Instruct, as of June 20, 2026

Prefix-cache evidence depends on the workload

One shared prefix

A single shared prefix is the easy case that any block-level cache handles well. It is not the workload SGLang's RadixAttention is built for.

Cache stateHit rateMeasured throughputFull condition
Cache on0 percent999 output tokens per secondOne shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of June 20, 2026.
Cache off0 percent910 output tokens per secondOne shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of June 20, 2026.
Cache on50 percent1,559 output tokens per secondOne shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of June 20, 2026.
Cache off50 percent912 output tokens per secondOne shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of June 20, 2026.
Cache on90 percent2,855 output tokens per secondOne shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of June 20, 2026.
Cache off90 percent910 output tokens per secondOne shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of June 20, 2026.

The highest published hit rate compares vLLM, cache on, 90 percent hit rate at 2,855 output tokens per second with vLLM, cache on, zero hit rate at 999 output tokens per second, under One shared prefix of 2,048 tokens sent to a varied fraction of requests, from 0 to 90 percent, with the engine's prefix cache on and off, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, cache state per row. Hardware and precision were not published. As of June 20, 2026.. The render-derived ratio is 2.86x.derived at render time for vLLM from its own cache-on throughput at a 90 percent hit rate against its cache-on throughput at a zero hit rate, output tokens per second, mean of three timed repeats after warmup, cache state per row

Many distinct prefixes

RadixAttention may pull ahead with longer prefixes, deeper trees, or heavier eviction pressure than we tested. On our test it did not, and we are not going to claim otherwise.

Distinct prefixesMeasured throughputFull condition
82,677 output tokens per second256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. NVIDIA H100 80GB. Precision was not published. As of June 20, 2026.
322,277 output tokens per second256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. NVIDIA H100 80GB. Precision was not published. As of June 20, 2026.
1281,467 output tokens per second256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on, at concurrency 32. output tokens per second, mean of three timed repeats after warmup, both engines with prefix caching on. NVIDIA H100 80GB. Precision was not published. As of June 20, 2026.

One client drives all three engines with identical request streams, so the timing definitions are the same everywhere. Time to first token is the time to the first streamed token. Throughput is output tokens per second. Warm up, then three timed repeats per operating point, charted as the mean; the 95 percent confidence intervals are tight and live in the committed raw data. Prefix caching is off for the unique-prompt sweeps. Concurrency swept from 1 to 256 at 1,024 input and 256 output tokens, unique prompts, prefix caching off. Same weights and same request stream across engines, single GPU, tensor-parallel size 1. 4 published package rows retain their recorded engine version, concurrency, and verified date.

We report package measurements here, and each package remains subject to its listed license.

Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.

Published packages keep their own measured conditions

These rows come from the published catalog, not the comparison sweep. Each package retains its own model, engine version, concurrency, and verified date.

Published package evidence, with every measurement tied to its package engine version, concurrency, and verified date
ModelPublished factsP50 latencyP99 latencyThroughputRecorded conditions
AREX-TurboBAAI/AREX-Turbo
  • optimized build of BAAI/AREX-Turbo
  • Channelwise FP8 weights, dynamic per-token activations
  • H100 via vLLM
  • 1.2x faster
  • gsm8k no measurable accuracy change, passed
  • USD $15 one-time, license Apache-2.0
666 ms baseline
551 ms optimized
684 ms baseline
579 ms optimized
1534 tokens/s baseline
1847 tokens/s optimized
vLLM 0.25.1
Concurrency 8
Jul 27, 2026
1.2x faster
gsm8k no measurable accuracy change, passed
Channelwise FP8 weights, dynamic per-token activations
Kimi K3moonshotai/Kimi-K3
  • optimized build of moonshotai/Kimi-K3
  • Technique not published
  • B300 via vLLM
  • 2.12x faster
  • Parity by construction
  • USD $100 one-time, license Kimi K3 License
19090 ms baseline
8978 ms optimized
19377 ms baseline
9595 ms optimized
55.2 tokens/s baseline
120 tokens/s optimized
vLLM 0.23.1
Concurrency 1
Jul 31, 2026
2.12x faster
Parity by construction
Not published
Qwen3.6 27BQwen/Qwen3.6-27B
  • optimized build of Qwen/Qwen3.6-27B
  • Channelwise FP8, measured selective-layer recipe
  • H100 via vLLM
  • 1.28x faster
  • gsm8k 99.87% recovery, passed
  • USD $40 one-time, license Apache-2.0
2857 ms baseline
2215 ms optimized
2881 ms baseline
2230 ms optimized
361 tokens/s baseline
465 tokens/s optimized
vLLM 0.25.1
Concurrency 8
Jul 25, 2026
1.28x faster
gsm8k 99.87% recovery, passed
Channelwise FP8, measured selective-layer recipe
Qwythos-9B-Claude-Mythos-5-1Mempero-ai/Qwythos-9B-Claude-Mythos-5-1M
  • optimized build of empero-ai/Qwythos-9B-Claude-Mythos-5-1M
  • Channelwise FP8 weights, dynamic per-token activations
  • H100 via vLLM
  • 1.29x faster
  • gsm8k 99.35% recovery, passed
  • USD $20 one-time, license Apache-2.0
1058 ms baseline
815 ms optimized
1083 ms baseline
861 ms optimized
974 tokens/s baseline
1255 tokens/s optimized
vLLM 0.25.1
Concurrency 8
Jul 25, 2026
1.29x faster
gsm8k 99.35% recovery, passed
Channelwise FP8 weights, dynamic per-token activations

The limitations bound every result

  • We benchmarked TensorRT-LLM on its PyTorch backend, which has no ahead-of-time engine compile, so its peak throughput here is not the ceiling a compiled TensorRT engine would reach. And our single-prefix cache test is not the many-distinct-prefix case SGLang's RadixAttention is designed for, so we ran that separately too. We did not let either feed the headline.
  • One H100 80GB and one L40S 48GB, single GPU, tensor-parallel size 1.
  • Same Llama-3.1-8B-Instruct weights and same request stream across engines. Three timed repeats; the 95 percent confidence intervals are tight and live in the committed raw data.
  • Latency here is TTFT and per-request percentiles, not goodput at a fixed SLO. Whether 514 ms first-token at high load is acceptable depends on your use case; an SLO-goodput sweep is future work.
  • Numbers are self-reported on our harness. We do not cross-calibrate against an external suite like MLPerf, so the open harness is the check: re-run it and compare.
  • TensorRT-LLM ran the PyTorch backend (no engine compile), so there is no compile time to report and its peak number may differ from a built TensorRT engine.
  • So the 4,813 tok/s we measured is the PyTorch-backend peak, not TensorRT-LLM's ceiling, and we got no compile-time number.
  • The L40S sweep stopped at concurrency 128 (its saturation), the H100 at 256.
  • Engines move weekly. These numbers are vLLM 0.23.0, SGLang 0.5.13, TensorRT-LLM 1.2.1, as of June 20 2026.
  • A single shared prefix is the easy case that any block-level cache handles well. It is not the workload SGLang's RadixAttention is built for.
  • RadixAttention may pull ahead with longer prefixes, deeper trees, or heavier eviction pressure than we tested. On our test it did not, and we are not going to claim otherwise.

What is not measured

  • No other model is measured in this comparison dataset. It covers Llama 3.1 8B Instruct.
  • No GPU outside NVIDIA H100 80GB and NVIDIA L40S 48GB is represented in the dataset.
  • No engine version newer than vLLM 0.23.0 is measured for this page's comparison evidence.
  • No fp8 throughput and no fp8 latency. The source publishes fp8 only as a derived cost per 1M output tokens at saturation, so no fp8 tokens-per-second, first-token latency, or speedup may be rendered from this dataset.
  • No L40S throughput and no L40S latency. The L40S appears only as a derived cost cell. Every sweep row here is the H100 at bf16.
  • No TensorRT-LLM prefix-cache result. Both prefix experiments ran vLLM and SGLang only.
  • No compiled TensorRT engine. Every TensorRT-LLM number here is its PyTorch backend, which the source states is not TensorRT-LLM's throughput ceiling, and no engine compile time was recorded.
  • No prefix-cache latency values. The source states one engine had lower first-token latency on the many-prefix workload but publishes no latency numbers for either prefix experiment, so that comparison stays qualitative.
  • No hardware or precision label for the single-shared-prefix sweep. Its caption names only the concurrency, so those rows must not be presented under a GPU or a precision the source does not state.
  • No goodput at a fixed service-level objective, and no accuracy or output-quality comparison between engines. This dataset is throughput, first-token latency, and derived cost only.
  • No cross-engine or cross-precision ratio is stored. Pages derive every multiplier and percentage from the absolute pair at render time and print both absolute values beside it.
  • Not cross-calibrated against an external benchmark suite. These are self-reported numbers from an open harness, and the harness is the check.
  • No claim that these results still hold. The dataset is pinned to one as-of date and one engine version triple, and engines change on a weekly cadence.

The same evidence answers common engine questions

How did vLLM compare with SGLang on throughput at the highest measured concurrency?
We measured higher output throughput for vLLM. vLLM recorded 5,333 output tokens per second; SGLang recorded 5,235 output tokens per second. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026.
How did vLLM compare with TensorRT-LLM on TTFT p50 at the highest measured concurrency?
We measured lower TTFT p50 for vLLM. vLLM recorded 2,221 milliseconds; TensorRT-LLM recorded 2,848 milliseconds. Condition: NVIDIA H100 80GB, BF16, concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026.
How did vLLM compare with SGLang on derived cost in the shared published configuration?
We derived lower cost for vLLM. vLLM was $0.206; SGLang was $0.21 USD per 1M output tokens. Condition: NVIDIA H100 80GB, BF16, source-selected cost concurrency 256, Llama 3.1 8B Instruct, as of June 20, 2026. Both values are derived from the recorded GPU rate and measured saturation throughput.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy