RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy

Optimization technique

Channelwise FP8 weights, dynamic per-token activations: measured speed and accuracy on open models

Channelwise FP8 stores weights in eight-bit floating point and scales activations for each token at runtime.

We measured 2 published packages across 1 GPU target, using the recorded baseline and optimized conditions.

What made these builds shippable is not the FP8 itself but the measured accuracy verdict beside it, taken on the same protocol as the baseline.

The published pairs set the limit

We show absolute baseline and optimized values before the shared verdict. We keep each condition attached to its result.

Published baseline and optimized runs at their recorded profiles, concurrency, and benchmark basis
ModelTargetThroughputMedian latencyVerdict and conditionsAccuracyPrice and date
AREX-TurboBAAI/AREX-TurboH100vLLM 0.25.11534 tokens/s baseline1847 tokens/s optimized666 ms p50 baseline551 ms p50 optimized1.2x faster.baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis.Baseline 0.3821, optimized 0.4094.gsm8k no measurable accuracy change, passedStandard errors: baseline +/- 0.0134; optimized +/- 0.0135.$15Verified Jul 27, 2026
Qwythos-9B-Claude-Mythos-5-1Mempero-ai/Qwythos-9B-Claude-Mythos-5-1MH100vLLM 0.25.1974 tokens/s baseline1255 tokens/s optimized1058 ms p50 baseline815 ms p50 optimized1.29x faster.baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis.Baseline 0.8378, optimized 0.8324.gsm8k 99.35% recovery, passedStandard errors not published.$20Verified Jul 25, 2026

Throughput changes by model

We restate the published throughput pairs. The neutral fills do not declare a winner.

AREX-Turbo

Baseline1534 tokens/s
Optimized1847 tokens/s

1.2x faster.

Qwythos-9B-Claude-Mythos-5-1M

Baseline974 tokens/s
Optimized1255 tokens/s

1.29x faster.

Output tokens per second under each package's recorded request profile and concurrency

Accuracy, honestly

We do not call a score change an improvement when the gap is smaller than about twice the combined standard errors.

AREX-Turbo measured 0.3821 at baseline and 0.4094 optimized on gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2). No measurable change in accuracy. The difference between the baseline and optimized scores is smaller than the combined measurement error, so it is within measurement noise. Standard errors: baseline +/- 0.0134; optimized +/- 0.0135.

Qwythos-9B-Claude-Mythos-5-1M measured 0.8378 at baseline and 0.8324 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. The optimized score measured below the baseline, at 99.35% of it. Neither score published a standard error, so this difference cannot be separated from sampling noise. The difference is not stated as a regression. Standard errors not published.

What is not measured

Tasks beyond gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2) and gsm8k exact_match strict (N=1319, completion protocol) not published.

GPUs beyond H100 not published.

Serving engines beyond vLLM 0.25.1 not published.

Questions buyers ask

Does using Channelwise FP8 weights, dynamic per-token activations hurt accuracy on these models?+

We tie each accuracy verdict to its published task metric. AREX-Turbo: Throughput measured 1534 tokens/s at baseline and 1847 tokens/s optimized. P50 latency measured 666 ms at baseline and 551 ms optimized. 1.2x faster. AREX-Turbo measured 0.3821 at baseline and 0.4094 optimized on gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2). No measurable change in accuracy. Standard errors: baseline +/- 0.0134; optimized +/- 0.0135. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 27, 2026. Qwythos-9B-Claude-Mythos-5-1M: Throughput measured 974 tokens/s at baseline and 1255 tokens/s optimized. P50 latency measured 1058 ms at baseline and 815 ms optimized. 1.29x faster. Qwythos-9B-Claude-Mythos-5-1M measured 0.8378 at baseline and 0.8324 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026.

What does using Channelwise FP8 weights, dynamic per-token activations buy on H100?+

We compare throughput and p50 latency under the recorded conditions. AREX-Turbo: Throughput measured 1534 tokens/s at baseline and 1847 tokens/s optimized. P50 latency measured 666 ms at baseline and 551 ms optimized. 1.2x faster. AREX-Turbo measured 0.3821 at baseline and 0.4094 optimized on gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2). No measurable change in accuracy. Standard errors: baseline +/- 0.0134; optimized +/- 0.0135. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 27, 2026. Qwythos-9B-Claude-Mythos-5-1M: Throughput measured 974 tokens/s at baseline and 1255 tokens/s optimized. P50 latency measured 1058 ms at baseline and 815 ms optimized. 1.29x faster. Qwythos-9B-Claude-Mythos-5-1M measured 0.8378 at baseline and 0.8324 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026.

Which open models were measured with Channelwise FP8 weights, dynamic per-token activations?+

AREX-Turbo and Qwythos-9B-Claude-Mythos-5-1M are the published models on this page. AREX-Turbo: Throughput measured 1534 tokens/s at baseline and 1847 tokens/s optimized. P50 latency measured 666 ms at baseline and 551 ms optimized. 1.2x faster. AREX-Turbo measured 0.3821 at baseline and 0.4094 optimized on gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2). No measurable change in accuracy. Standard errors: baseline +/- 0.0134; optimized +/- 0.0135. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 27, 2026. Qwythos-9B-Claude-Mythos-5-1M: Throughput measured 974 tokens/s at baseline and 1255 tokens/s optimized. P50 latency measured 1058 ms at baseline and 815 ms optimized. 1.29x faster. Qwythos-9B-Claude-Mythos-5-1M measured 0.8378 at baseline and 0.8324 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026.

What remains unmeasured for Channelwise FP8 weights, dynamic per-token activations?+

AREX-Turbo: Throughput measured 1534 tokens/s at baseline and 1847 tokens/s optimized. P50 latency measured 666 ms at baseline and 551 ms optimized. 1.2x faster. AREX-Turbo measured 0.3821 at baseline and 0.4094 optimized on gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2). No measurable change in accuracy. Standard errors: baseline +/- 0.0134; optimized +/- 0.0135. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 27, 2026. Qwythos-9B-Claude-Mythos-5-1M: Throughput measured 974 tokens/s at baseline and 1255 tokens/s optimized. P50 latency measured 1058 ms at baseline and 815 ms optimized. 1.29x faster. Qwythos-9B-Claude-Mythos-5-1M measured 0.8378 at baseline and 0.8324 optimized on gsm8k exact_match strict (N=1319, completion protocol). Measured difference, significance not published. Standard errors not published. baseline used request profile chat-128, concurrency 8, and measured basis; optimized used request profile chat-128, concurrency 8, and measured basis. Verified Jul 25, 2026. Tasks beyond gsm8k exact_match (N=1319 full set, chat-template, deliberation-aware extraction v2) and gsm8k exact_match strict (N=1319, completion protocol) not published. GPUs beyond H100 not published. Serving engines beyond vLLM 0.25.1 not published.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy