RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Weight-only versus weight-activation quantization

Quantization

Weight-only versus weight-activation quantization

What it is

Weight-only quantization narrows stored weights while activations remain at higher precision. Weight-activation quantization also narrows activations so supported hardware can use lower-precision compute paths.

Why it moves cost and latency

Weight-only methods primarily reduce model memory and weight bandwidth with simpler activation handling. Narrowing activations can reduce more traffic and use different compute units, but scaling and outliers become harder.

What it looks like in practice

Teams compare both on the same model, evaluation set, request shapes, and hardware. Kernel availability, memory limits, quality tolerance, and workload shape decide the better configuration.

Related terms

  • Channelwise quantization->
  • Four-bit weight-only quantization->
  • Eight-bit integer quantization->
  • FP8 versus eight-bit integer quantization->
  • Mixed-precision serving->

Questions this definition answers

What does Weight-only versus weight-activation quantization mean in inference serving?

Weight-only quantization narrows stored weights while activations remain at higher precision. Weight-activation quantization also narrows activations so supported hardware can use lower-precision compute paths. Teams compare both on the same model, evaluation set, request shapes, and hardware. Kernel availability, memory limits, quality tolerance, and workload shape decide the better configuration.

Why can Weight-only versus weight-activation quantization move cost or latency?

Weight-only methods primarily reduce model memory and weight bandwidth with simpler activation handling. Narrowing activations can reduce more traffic and use different compute units, but scaling and outliers become harder.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy