RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/FP8 quantization

Quantization

FP8 quantization

What it is

FP8 quantization represents selected weights, activations, or cache values with eight-bit floating point formats. Scale factors map the narrower format to the magnitude range needed by each tensor or channel.

Why it moves cost and latency

Narrower values reduce memory traffic and storage, and supported hardware can execute more low-precision work. Poor scaling or sensitive layers can erase those gains through quality loss or fallback work.

What it looks like in practice

A serving stack chooses which tensors use the narrow format and how scales are produced. Teams compare task scores and serving metrics against the unquantized baseline before accepting the configuration.

Where we measured it

  • Measured dynamic channel FP8 results->
  • Measured staged dynamic channel FP8 results->
  • Measured H100 FP8 packages->
  • FP8 benchmark receipts and quality evidence->

Related terms

  • Channelwise quantization->
  • Tokens per second->
  • Tensor parallelism->
  • Quantization quality recovery->
  • SafeTensors->
  • Four-bit weight-only quantization->
  • Eight-bit integer quantization->
  • FP8 versus eight-bit integer quantization->
  • Mixed-precision serving->

Questions this definition answers

What does FP8 quantization mean in inference serving?

FP8 quantization represents selected weights, activations, or cache values with eight-bit floating point formats. Scale factors map the narrower format to the magnitude range needed by each tensor or channel. A serving stack chooses which tensors use the narrow format and how scales are produced. Teams compare task scores and serving metrics against the unquantized baseline before accepting the configuration.

Why can FP8 quantization move cost or latency?

Narrower values reduce memory traffic and storage, and supported hardware can execute more low-precision work. Poor scaling or sensitive layers can erase those gains through quality loss or fallback work.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy