RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Catalog/AREX-Turbo
Model

AREX-Turbo measured packages and prices

RunInfra has published 1 measured package for BAAI/AREX-Turbo on H100. Every package below keeps its engine, price, and proof date together.

Parameters
4.4B
Modality
Text
Base model license
Apache-2.0
Hugging Face model
BAAI/AREX-Turbo

The published packages are available to buy

Published packages for AREX-Turbo, with GPU, engine, optimization, price, and proof date
PackageGPUEngineOptimizationPriceVerified
arex-turbo-fp8cd-h100-vllmH100vLLM 0.25.1Channelwise FP8 weights, dynamic per-token activations$15 one-timeJul 27, 2026

What we measured

Channelwise FP8 changes the weights without reading calibration data, while activation scales are chosen per token at runtime.

arex-turbo-fp8cd-h100-vllm

Published Jul 27, 2026. Proof verified Jul 27, 2026.

1.2x faster

Median latency measured 666 ms at baseline and 551 ms optimized.

Quality evidence

gsm8k no measurable accuracy change, passed.

Published package measurements and prices for AREX-Turbo, grouped from the public RunInfra catalog.

Measured by RunInfra on H100 with vLLM 0.25.1; proof dates are listed with each package.

We report package measurements here, and each package remains subject to its listed license.

Citation: RunInfra (2026). Measured open-model serving benchmarks. https://runinfra.ai/catalog.

What is not measured

Other GPU measurements for this model not published.

Other engine measurements for this model not published.

Measurements for context lengths not listed in these packages not published.

Published facts answer common questions

What does AREX-Turbo cost on RunInfra?

arex-turbo-fp8cd-h100-vllm is listed at $15 one-time, published Jul 27, 2026.

Which GPU serves AREX-Turbo?

arex-turbo-fp8cd-h100-vllm is measured on H100 with vLLM 0.25.1. Proof verified Jul 27, 2026.

Is quality measured for AREX-Turbo?

gsm8k no measurable accuracy change, passed. Proof verified Jul 27, 2026.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy