RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Compare

Measured serving engine comparisons

The comparison dataset currently carries 3 measured pairs across 3 engines.

  1. vLLM vs SGLang

    Model Llama 3.1 8B Instruct; GPU set NVIDIA H100 80GB and NVIDIA L40S 48GB; engine versions vLLM 0.23.0 and SGLang 0.5.13; as of Jun 20, 2026.

  2. vLLM vs TensorRT-LLM

    Model Llama 3.1 8B Instruct; GPU set NVIDIA H100 80GB and NVIDIA L40S 48GB; engine versions vLLM 0.23.0 and TensorRT-LLM 1.2.1; as of Jun 20, 2026.

  3. SGLang vs TensorRT-LLM

    Model Llama 3.1 8B Instruct; GPU set NVIDIA H100 80GB and NVIDIA L40S 48GB; engine versions SGLang 0.5.13 and TensorRT-LLM 1.2.1; as of Jun 20, 2026.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy