RunInfraby RightNow
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Quantization

Glossary category

Quantization glossary

Definitions for inference serving terms in the quantization category.

Terms in this category

  • FP8 quantizationFP8 quantization represents selected weights, activations, or cache values with eight-bit floating point formats.
  • Channelwise quantizationChannelwise quantization assigns a separate scale to each channel instead of sharing one scale across an entire tensor.
  • Quantization quality recoveryQuantization quality recovery is the measured retention of task performance after a model moves to lower precision.
  • KV cache quantizationKV cache quantization stores cached attention keys and values in narrower numeric formats.
  • AWQActivation-aware weight quantization uses activation statistics to identify weight channels that need extra protection.
  • GPTQGPTQ is a layerwise post-training weight quantization method that uses approximate second-order information.
  • Four-bit weight-only quantizationFour-bit weight-only quantization stores model weights as four-bit integers while keeping activations at higher precision.
  • Eight-bit integer quantizationEight-bit integer quantization represents selected tensors with integer values and scale factors.
  • Weight-only versus weight-activation quantizationWeight-only quantization narrows stored weights while activations remain at higher precision.
  • FP8 versus eight-bit integer quantizationFP8 uses an exponent and mantissa within an eight-bit floating-point representation.
  • Calibration dataCalibration data is a sample of model inputs used to derive quantization scales or protection choices.
  • Outlier channelsOutlier channels contain values whose magnitudes are much larger than most peer channels.
  • Mixed-precision servingMixed-precision serving assigns different numeric precisions to tensors, operations, or layers in one model configuration.

Related glossary categories

  • Batching glossary->
  • Memory glossary->
  • Latency glossary->
  • Parallelism glossary->
  • Serving glossary->
  • Hardware glossary->
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy