RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/FP8 versus eight-bit integer quantization

Quantization

FP8 versus eight-bit integer quantization

What it is

FP8 uses an exponent and mantissa within an eight-bit floating-point representation. Eight-bit integer quantization uses uniformly spaced integer levels interpreted through scale factors.

Why it moves cost and latency

Floating point carries wider dynamic range, while integer levels can provide finer spacing inside a chosen range. Hardware support, tensor distributions, calibration burden, and quality targets decide the fit.

What it looks like in practice

Teams compare supported kernels on identical model tensors, workloads, and quality evaluations. They inspect outliers, fallback operations, scale granularity, memory traffic, and end-to-end latency.

Where we measured it

  • Measured dynamic FP8 scaling results->
  • Measured staged FP8 scaling results->
  • Measured H100 FP8 packages->

Related terms

  • FP8 quantization->
  • Eight-bit integer quantization->
  • Weight-only versus weight-activation quantization->
  • Calibration data->
  • Mixed-precision serving->

Questions this definition answers

What does FP8 versus eight-bit integer quantization mean in inference serving?

FP8 uses an exponent and mantissa within an eight-bit floating-point representation. Eight-bit integer quantization uses uniformly spaced integer levels interpreted through scale factors. Teams compare supported kernels on identical model tensors, workloads, and quality evaluations. They inspect outliers, fallback operations, scale granularity, memory traffic, and end-to-end latency.

Why can FP8 versus eight-bit integer quantization move cost or latency?

Floating point carries wider dynamic range, while integer levels can provide finer spacing inside a chosen range. Hardware support, tensor distributions, calibration burden, and quality targets decide the fit.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy