RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/GPTQ

Quantization

GPTQ

What it is

GPTQ is a layerwise post-training weight quantization method that uses approximate second-order information. It compensates remaining weights as quantization decisions are applied to limit layer output error.

Why it moves cost and latency

The method can produce compact weight-only checkpoints without retraining the full model. Quality and serving value depend on calibration data, quantization settings, packing, and runtime kernels.

What it looks like in practice

A quantizer collects layer inputs, processes weight groups, and writes a packed checkpoint. Teams evaluate task quality and serving performance on the exact exported artifact.

Related terms

  • Quantization quality recovery->
  • Calibration data->
  • Four-bit weight-only quantization->
  • AWQ->

Questions this definition answers

What does GPTQ mean in inference serving?

GPTQ is a layerwise post-training weight quantization method that uses approximate second-order information. It compensates remaining weights as quantization decisions are applied to limit layer output error. A quantizer collects layer inputs, processes weight groups, and writes a packed checkpoint. Teams evaluate task quality and serving performance on the exact exported artifact.

Why can GPTQ move cost or latency?

The method can produce compact weight-only checkpoints without retraining the full model. Quality and serving value depend on calibration data, quantization settings, packing, and runtime kernels.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy