RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Four-bit weight-only quantization

Quantization

Four-bit weight-only quantization

What it is

Four-bit weight-only quantization stores model weights as four-bit integers while keeping activations at higher precision. Serving kernels dequantize or rescale weights as they participate in matrix operations.

Why it moves cost and latency

Narrow weights reduce checkpoint size, device memory use, and weight bandwidth. Dequantization overhead, packing support, and quantization error decide whether those savings improve serving.

What it looks like in practice

Weights are grouped, scaled, packed, and loaded by kernels that understand the layout. Teams compare memory, throughput, latency, and task quality against a wider baseline.

Related terms

  • FP8 quantization->
  • AWQ->
  • GPTQ->
  • GGUF->
  • Weight-only versus weight-activation quantization->

Questions this definition answers

What does Four-bit weight-only quantization mean in inference serving?

Four-bit weight-only quantization stores model weights as four-bit integers while keeping activations at higher precision. Serving kernels dequantize or rescale weights as they participate in matrix operations. Weights are grouped, scaled, packed, and loaded by kernels that understand the layout. Teams compare memory, throughput, latency, and task quality against a wider baseline.

Why can Four-bit weight-only quantization move cost or latency?

Narrow weights reduce checkpoint size, device memory use, and weight bandwidth. Dequantization overhead, packing support, and quantization error decide whether those savings improve serving.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy