RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Mixed-precision serving

Quantization

Mixed-precision serving

What it is

Mixed-precision serving assigns different numeric precisions to tensors, operations, or layers in one model configuration. Sensitive paths retain wider formats while less sensitive paths use narrower ones.

Why it moves cost and latency

Selective precision can recover quality while keeping much of the memory or compute benefit. Format conversions, fallback kernels, and poor placement can offset the intended gain.

What it looks like in practice

Teams vary precision by component, then measure quality and serving behavior against a consistent baseline. The accepted map records every wider fallback so results remain reproducible.

Where we measured it

  • Measured mixed-precision scaling results->
  • Measured staged mixed-precision results->
  • Measured H100 mixed-precision packages->

Related terms

  • FP8 quantization->
  • KV cache quantization->
  • Weight-only versus weight-activation quantization->
  • FP8 versus eight-bit integer quantization->
  • Outlier channels->

Questions this definition answers

What does Mixed-precision serving mean in inference serving?

Mixed-precision serving assigns different numeric precisions to tensors, operations, or layers in one model configuration. Sensitive paths retain wider formats while less sensitive paths use narrower ones. Teams vary precision by component, then measure quality and serving behavior against a consistent baseline. The accepted map records every wider fallback so results remain reproducible.

Why can Mixed-precision serving move cost or latency?

Selective precision can recover quality while keeping much of the memory or compute benefit. Format conversions, fallback kernels, and poor placement can offset the intended gain.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy