RunInfraby RightNow
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Engines

Measured serving engines for open model inference

The published evidence currently covers 3 engines and 4 published packages.

  1. SGLang

    0 published packages

    Comparison scope
    Llama 3.1 8B Instruct on NVIDIA H100 80GB and NVIDIA L40S 48GB, measured Jun 20, 2026.
    Measured versions
    0.5.13
    Evidence sources
    Comparison rows and published package records stay separated on the detail page.
  2. TensorRT-LLM

    0 published packages

    Comparison scope
    Llama 3.1 8B Instruct on NVIDIA H100 80GB and NVIDIA L40S 48GB, measured Jun 20, 2026.
    Measured versions
    1.2.1
    Evidence sources
    Comparison rows and published package records stay separated on the detail page.
  3. vLLM

    4 published packages

    Comparison scope
    Llama 3.1 8B Instruct on NVIDIA H100 80GB and NVIDIA L40S 48GB, measured Jun 20, 2026.
    Measured versions
    0.23.0, 0.23.1, and 0.25.1
    Evidence sources
    Comparison rows and published package records stay separated on the detail page.
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy