arex-turbo-fp8cd-h100-vllm
Published Jul 27, 2026. Proof verified Jul 27, 2026.
1.2x faster
Median latency measured 666 ms at baseline and 551 ms optimized.
Quality evidence
gsm8k no measurable accuracy change, passed.
RunInfra has published 1 measured package for BAAI/AREX-Turbo on H100. Every package below keeps its engine, price, and proof date together.
| Package | GPU | Engine | Optimization | Price | Verified |
|---|---|---|---|---|---|
| arex-turbo-fp8cd-h100-vllm | H100 | vLLM 0.25.1 | Channelwise FP8 weights, dynamic per-token activations | $15 one-time | Jul 27, 2026 |
Channelwise FP8 changes the weights without reading calibration data, while activation scales are chosen per token at runtime.
Published Jul 27, 2026. Proof verified Jul 27, 2026.
Median latency measured 666 ms at baseline and 551 ms optimized.
gsm8k no measurable accuracy change, passed.
Published package measurements and prices for AREX-Turbo, grouped from the public RunInfra catalog.
Measured by RunInfra on H100 with vLLM 0.25.1; proof dates are listed with each package.
We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.
Other GPU measurements for this model not published.
Other engine measurements for this model not published.
Measurements for context lengths not listed in these packages not published.
arex-turbo-fp8cd-h100-vllm is listed at $15 one-time, published Jul 27, 2026.
arex-turbo-fp8cd-h100-vllm is measured on H100 with vLLM 0.25.1. Proof verified Jul 27, 2026.
gsm8k no measurable accuracy change, passed. Proof verified Jul 27, 2026.
© 2026 RunInfra. All rights reserved.