RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading the optimized model catalog

Optimized models ready for production serving

Open models, quantized and benchmarked on real GPUs. Buy once, run anywhere.

Loading the model package
← All optimized models

Qwen3-1.7B

Qwen/Qwen3-1.7B

Quantized and serving-tuned for H100 on vLLM 0.25.1, with the benchmark receipt included. You get a self-contained kit: the optimized weights, the exact serving configuration, and the measured proof. Run it on any cloud or your own hardware. Nothing calls home, nothing expires.

$10
one-time, yours forever
+ $20/mo continuous optimization
Get this model
$20/mo

We keep re-optimizing and re-benchmarking as engines move, and you always download the latest verified version. Cancel anytime. Nothing is charged until you confirm at checkout.

On. Confirmed at checkout.

Latency p50
294ms, concurrency 8
Latency p99
307ms, concurrency 8
Accuracy
100%
gsm8k exact_match strict (N=100, chat-template reasoning-aware harness, extraction v2) recoveryPassed accuracy gate
GPU target
H100
Base model license
Apache-2.0
Proof verified
Jul 27, 2026
Published
Jul 27, 2026

Company

Who you are buying from

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
  • Y CombinatorBacked by Y Combinator.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
Zero data retention
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

All systems operational
Pipeline BuilderModelsPricingStartupsDocs
Latency p50, Concurrency 8lower is better↓ 9%
Baseline322 ms
Optimized294 ms
0250500
Latency p99, Concurrency 8lower is better↓ 7%
Baseline328 ms
Optimized307 ms
0250500
Throughput, Concurrency 8higher is better↑ 9%
Baseline3163 tokens/s
Optimized3463 tokens/s
025005000

Runs on any GPU provider

The kit is self-contained: bring it to any cloud with H100-class GPUs, or your own hardware.

Docker Compose
Kubernetes
RunPod
Modal
Your own hardware

This is what it means to own your AI

You control the intelligence

The weights run on your infrastructure. No deprecation notices, no silent model swaps, no provider deciding what your product is allowed to do.

You keep the economics

One purchase instead of a meter that runs forever. A closed API charges you for every token, at whatever the price becomes; your own deployment costs what your GPU costs.

You keep the data

Prompts and outputs never leave your network. Nothing calls home, nothing is retained by someone else, and compliance stops being a negotiation.

Technical implementation

Channelwise FP8 weights, dynamic per-token activations

llm-compressor 0.12.0, data-free: no calibration corpus is read at any point, so the quantization cannot have seen evaluation data

Serving configuration swept per model and pinned

the exact engine build and arguments every receipt number was produced on: vLLM 0.25.1, max_num_seqs 256, max_num_batched_tokens 8192, max-model-len 4096, chosen as the best of a five-point grid measured on this model rather than inherited from another package

Serving stack

vLLM 0.25.1, Channelwise FP8 weights, dynamic per-token activations, tuned for H100.

Fit-tested minimum
Not published
Published GPU capacity is not a proven minimum fit requirement.

Measured accuracy

gsm8k exact_match strict

(N=100, chat-template reasoning-aware harness, extraction v2)

Passed accuracy gate

Baseline
0.8±0.04 standard error
Optimized
0.8±0.04 standard error
Recovery
100%
Delta
0%

No measurable change in accuracy

The difference between the baseline and optimized scores is smaller than the combined measurement error, so it is within measurement noise.

Methodology

  • NO MEASURABLE CHANGE: 80/100 on both sides, identical at the precision this instrument resolves, with a standard error of 0.04 per side. That is 100.0 percent recovery, equivalently a 0.0 percent relative capability difference. Read it as indistinguishable rather than as proof of exact equality: with N=100 this measurement could not detect a difference smaller than roughly 8 points. Zero truncation and zero invalid terminals on the shipped-artifact side. lm-eval-class harness with the chat template applied, seed 1234, max_tokens 10000, temperature 0, streamed, non-clean terminals scored wrong, identical settings and the same H100 class on both sides, with the optimized side bound to the shipped artifact sha256 233a04cb. The gate verdict covers this gsm8k metric;
  • ifeval, tool calling and safety refusal delta are recorded separately in accuracy-merged.json and all three returned a gate-eligible state with no exemption.

Buyer disclosure

What we did not measure

  • Refusal, toxicity and jailbreak behavior: the refusal-delta suite measured the CHANGE quantization caused against this model's own BF16 baseline and found it within tolerance. Absolute safety was not certified, and toxicity and jailbreak resistance were not evaluated. Run your own safety evaluation before user-facing or regulated deployment.
  • Time to first token is slower on the optimized build: 29.8 ms against the baseline's 23.4 ms at the median. End-to-end latency is still better (293.8 ms against 322.3 ms p50) because decode more than repays it, but if your workload is dominated by very short completions, weigh the first-token cost.
  • The speed win is modest and measured, not the figure FP8 is often assumed to deliver: about 1.09x output tokens per second and 1.10x on end-to-end p50, at concurrency 8 on the chat-128 profile. A custom fused SwiGLU kernel was built and measured for this model and rejected as an end-to-end wash, proven by an engagement counter in the receipt rather than assumed, because vLLM's stock merged-GEMM plus SiluAndMul is already near optimal at decode tile sizes.

License

Two licenses apply: the RunInfra package license you purchase under, and the base model's own open-source license.

Package license

RunInfra Package License v1.1 (2026-07-25)

Buy once. No meter, no expiry, nothing calling home. The full terms ship inside the kit.

What you are licensed to do

This package is licensed to the purchasing workspace, one time, for the exact version purchased. Everyone in that workspace may run it in production without limits: unlimited inference, on any hardware or cloud the workspace controls, for any lawful commercial purpose. You may modify the configuration and tooling for your own use.

Your outputs are yours

Everything the model produces for you belongs to you. You may use, sell, and build products on the model's outputs without restriction or royalty. Serving the model to your own customers as part of your product is use, not redistribution, and is fully allowed.

What you may not do

You may not redistribute, resell, sublicense, rent, publish, or otherwise make the package or its artifacts available to any third party. That covers the optimized weights, serving configuration, scripts, kit archive, and benchmark receipts, whole or in part, modified or not. The license belongs to the purchasing workspace and cannot be transferred separately from it.

The base model keeps its own license

The underlying model remains under its upstream open-source license, which is included in this kit with attribution and a statement of RunInfra's modifications. Nothing in this license restricts rights the upstream license grants you for the ORIGINAL model; the restrictions above apply to RunInfra's optimized package.

Version-pinned, as measured

You purchased this exact version, proven against the exact engine version named in the receipt. It stays downloadable to your workspace and never expires. The benchmark receipt describes measurements taken at verification time on the named hardware; RunInfra does not promise future updates to this version, and later package versions are separate purchases.

Breach ends the license

If the workspace redistributes the package or its artifacts, this license terminates for that workspace. Sections about your outputs survive termination for outputs already produced.

Base model license

Apache-2.0, as published by the model author on Hugging Face. The full license text ships inside the kit.

  • NVIDIA InceptionMember of NVIDIA Inception.
  • Research
    News
    Contact
    Backed by
    YCombinator
    AICPA Type II
    SOC 2
    NVIDIA Inception ProgramNVIDIA Inception Program
    Ask AI about RunInfra
    Part of RightNow
    SecurityDPAAUPCookiesTermsPrivacy
  • Capability evidence is gsm8k at N=100, not the full 1319-item set, plus the three-suite battery. At that sample size the measurement resolves differences of roughly 8 points, so it establishes the absence of a large regression rather than exact parity.
  • Minimum VRAM is listed as 80 GB because that is what the shipped kit manifest records and the kit is what your verifier checks against. It describes the card this package was measured on, not the model's requirement: the FP8 weights are 2.15 GB. The package has only been characterized on H100, so smaller cards are untested rather than unsupported.
  • Measured at 4096 context, concurrency 8, on a single H100 with a per-model swept serving configuration (max_num_seqs 256, max_num_batched_tokens 8192). Longer context, other GPUs and higher concurrency are not characterized, though a soak at concurrency 48 completed 256 of 256 requests with zero failures at 12837 output tokens per second.
  • This is a hybrid thinking model measured in its default serve mode. Thinking-mode-specific behavior and long-context behavior are not separately certified.
  • Calibration
    No contamination risk: fp8_channel_dynamic is DATA-FREE (round-to-nearest weights plus dynamic per-token activation scales). No calibration corpus is read at any point, so overlap between calibration data and evaluation data is structurally impossible rather than merely unlikely.
    Cost basis
    Cost inputs current as of 2026-07-27; economics at the measured operating point: 3462.8 output tokens/sec on one H100 at the chat-128 profile, concurrency 8, which at an assumed USD 4.50/hr is about USD 0.36 per 1M output tokens. At soak-48 saturation (12837 tokens/sec) it is about USD 0.10 per 1M. Throughput is measured; the hourly rate is an assumption, so recompute for your own rate.