vLLM vs SGLang
Measured sweep Llama 3.1 8B Instruct on NVIDIA H100 80GB at BF16; derived cost configurations H100 BF16, H100 FP8, and L40S BF16; engine versions vLLM 0.23.0 and SGLang 0.5.13; as of Jun 20, 2026.
Compare
The comparison dataset currently carries 3 measured pairs across 3 engines.
Measured sweep Llama 3.1 8B Instruct on NVIDIA H100 80GB at BF16; derived cost configurations H100 BF16, H100 FP8, and L40S BF16; engine versions vLLM 0.23.0 and SGLang 0.5.13; as of Jun 20, 2026.
Measured sweep Llama 3.1 8B Instruct on NVIDIA H100 80GB at BF16; derived cost configurations H100 BF16, H100 FP8, and L40S BF16; engine versions vLLM 0.23.0 and TensorRT-LLM 1.2.1; as of Jun 20, 2026.
Measured sweep Llama 3.1 8B Instruct on NVIDIA H100 80GB at BF16; derived cost configurations H100 BF16, H100 FP8, and L40S BF16; engine versions SGLang 0.5.13 and TensorRT-LLM 1.2.1; as of Jun 20, 2026.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs