Skip to main content
The Optimized Models Catalog contains verified optimized model packages for open-weight models. Each package combines one model, one target GPU, and one serving engine. RunInfra quantizes and serving-tunes that combination, then publishes its measured speedup, throughput, accuracy recovery, and VRAM receipt.
Measurements apply to the model, GPU, and serving engine named on the package page. Review that page before buying instead of assuming the same results for another configuration.

What a package includes

Browse the catalog

Browse the public catalog at runinfra.ai/catalog. Each model has a public page at /catalog/<slug> with:
  • Measured proof charts
  • Package specifications, including the target GPU and serving engine
  • A provider comparison
  • Package and base model license information
  • A shareable model card
The measured figures vary by package. Use the figures on the model page as the source of truth for that package.

Buy a package

1

Review the package

Confirm the model, target GPU, serving engine, measured receipt, and both licenses on the model page.
2

Check your workspace credits

Packages use your existing workspace credit balance, where 1 credit = $1. Your workspace must have a paid plan before it can spend credits on a package.
3

Complete the one-time purchase

The package belongs to the purchasing workspace. It never expires, and you can download it again at any time.
Continuous Optimization is an optional monthly subscription. It is separate from the one-time package purchase.

Download the kit

After purchase, download the self-contained kit from the package page. It contains the optimized weights, exact serving configuration, measured proof, license files, and deployment guides. The kit does not call home. You can run it on the package’s target GPU in any cloud or on your own hardware.

Deploy the package

Choose the included guide for your environment: For the Docker Compose quickstart, unzip the kit and open its directory. Then set VLLM_API_KEY and start the included Compose configuration:
After the service starts, use the exact endpoint from the included guide:
Keep the model, GPU, serving engine, and serving configuration together. The receipt describes that measured combination.

Understand the licenses

Two licenses apply to every package: You must follow both licenses. Buying a package does not replace or override the base model’s upstream terms.

Request a model

If the catalog does not include the combination you need, submit a request with:
  • Model
  • GPU
  • Target
  • Serving engine
RunInfra uses request volume to decide which combinations to optimize next. The most-requested combinations move to the front of the catalog queue.

Catalog packages and custom optimization

The catalog gives you a prebuilt package for a listed model, GPU, and serving engine. Use Optimization when you want RunInfra to measure and rank variants for your own pipeline and constraints. See Models for the broader set of models RunInfra can resolve and deploy.

Next steps

Continuous Optimization

Keep a purchased package current with newly verified versions.

Models

Browse the model types RunInfra supports.

Optimization

Measure and rank variants for your own pipeline.

Plans and pricing

Learn how workspace credits and paid plans work.