arex-turbo-fp8cd-h100-vllm
Published Jul 27, 2026. Proof verified Jul 27, 2026.
1.2x faster
Median latency measured 666 ms at baseline and 551 ms optimized.
Quality evidence
gsm8k no measurable accuracy change, passed.
RunInfra lists 1 measured package for AREX-Turbo. The newest published row, arex-turbo-fp8cd-h100-vllm, costs $15 one-time on H100 with vLLM 0.25.1. Its speed result is 1.2x faster. Request profile: chat-128. Concurrency 8. That result applies only to those recorded conditions. Proof verified Jul 27, 2026. Other package prices and measurements remain separate below.
| Package | GPU | Engine | Optimization | Price | Verified |
|---|---|---|---|---|---|
| arex-turbo-fp8cd-h100-vllm | H100 | vLLM 0.25.1 | Channelwise FP8 weights, dynamic per-token activations | $15 one-time | Jul 27, 2026 |
Channelwise FP8 changes the weights without reading calibration data, while activation scales are chosen per token at runtime.
Published Jul 27, 2026. Proof verified Jul 27, 2026.
Median latency measured 666 ms at baseline and 551 ms optimized.
gsm8k no measurable accuracy change, passed.
Published package measurements and prices for AREX-Turbo, grouped from the public RunInfra catalog.
Measured by RunInfra on H100 with vLLM 0.25.1; proof dates are listed with each package.
We report package measurements here, and each package remains subject to its listed license.
Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/benchmarks.
Other GPU measurements for this model not published.
Other engine measurements for this model not published.
Measurements for context lengths not listed in these packages not published.
arex-turbo-fp8cd-h100-vllm is listed at $15 one-time, published Jul 27, 2026.
arex-turbo-fp8cd-h100-vllm is measured on H100 with vLLM 0.25.1. Proof verified Jul 27, 2026.
gsm8k no measurable accuracy change, passed. Proof verified Jul 27, 2026.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs