What does FP8 versus eight-bit integer quantization mean in inference serving?
FP8 uses an exponent and mantissa within an eight-bit floating-point representation. Eight-bit integer quantization uses uniformly spaced integer levels interpreted through scale factors. Teams compare supported kernels on identical model tensors, workloads, and quality evaluations. They inspect outliers, fallback operations, scale granularity, memory traffic, and end-to-end latency.