Llama 3.1 8B Instruct
NVIDIA A100 (40 GB), measured 2026-07-12
- 71.12 ms
- Time to first token
- 391.76 tok/s
- Output throughput
- $1.489
- Per 1M output tokens
- 0.973
- Quality vs FP16 baseline
p50, batched serving (FP16, batch 24)
measured, batched serving (FP16, batch 24)
derived from recorded GPU rate and measured saturation, batched serving config
measured, FP8 winner config
Source record llama31-8b-a100-2026-07-12