# RunInfra (full text) ## Model APIs Model APIs is RunInfra's hosted model API product at https://api.runinfra.ai/v1. Each model publishes its operation and compatibility. Chat models expose OpenAI-compatible chat completions at POST /v1/chat/completions and Anthropic-compatible Messages at POST /v1/messages. Chat model requests use workspace API keys and support streaming. JSON mode is supported per chat model, and a model that requires a specific reasoning_effort to deliver it publishes that value with its capabilities. Pricing is published per model in USD with only the billed dimensions shown. This file also contains technical reference material, use-case recipes, GPU reference data, and newsroom articles selected for answer-engine ingestion without rendering the UI. The entries below come from the hosted model registry. A registered model appears here only when its own public row passes the same publication gate used by the sitemap. Entries that do not pass that public gate are omitted. - Model API library (https://runinfra.ai/inference-api): every published hosted model API is listed here with its operation and compatibility contract. - DeepSeek V4 Flash Model API (https://runinfra.ai/inference-api/deepseek-v4-flash): Served model ID: deepseek-ai/DeepSeek-V4-Flash-0731. OpenAI-compatible chat completions: POST /v1/chat/completions. Anthropic-compatible Messages: POST /v1/messages. Published token pricing: $0.13 per 1M input tokens; $0.01 per 1M cached input tokens; $0.27 per 1M output tokens. Context window: 1,048,576 tokens. Published capabilities: Tool calling, JSON mode, Streaming. Omitted reasoning_effort is sent as maximum. Requests that call tools, without response_format, are sent as medium. Requests that set response_format are sent as none. API access is unavailable. Next check at Sep 9, 2026, 10:27 PM UTC. - Nemotron 3.5 Lightning 30B Model API (https://runinfra.ai/inference-api/nemotron-3-5-lightning-30b): Served model ID: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16. OpenAI-compatible chat completions: POST /v1/chat/completions. Anthropic-compatible Messages: POST /v1/messages. Published token pricing: $0.05 per 1M input tokens; $0.01 per 1M cached input tokens; $0.15 per 1M output tokens. Context window: 262,144 tokens. Published capabilities: Tool calling, JSON mode, Streaming. API access is available. - Qwen3.8 27B Model API (https://runinfra.ai/inference-api/qwen3-8-27b): Served model ID: Qwen/Qwen3.8-27B. OpenAI-compatible chat completions: POST /v1/chat/completions. Anthropic-compatible Messages: POST /v1/messages. Published token pricing: $0.10 per 1M input tokens; $0.01 per 1M cached input tokens; $0.40 per 1M output tokens. Context window: 262,144 tokens. Published capabilities: Tool calling, JSON mode, Streaming. API access is available. - DeepSeek V4 Pro Model API (https://runinfra.ai/inference-api/deepseek-v4-pro): Served model ID: deepseek-ai/DeepSeek-V4-Pro-0813. OpenAI-compatible chat completions: POST /v1/chat/completions. Anthropic-compatible Messages: POST /v1/messages. Published token pricing: $0.60 per 1M input tokens; $0.03 per 1M cached input tokens; $1.90 per 1M output tokens. Context window: 1,048,576 tokens. Published capabilities: Tool calling, JSON mode, Streaming. API access is unavailable. Next check at Sep 9, 2026, 10:29 PM UTC. - Ornith 1.5 35B Model API (https://runinfra.ai/inference-api/ornith-1-5-35b): Served model ID: ornith-ai/Ornith-1.5-35B-A3B. OpenAI-compatible chat completions: POST /v1/chat/completions. Anthropic-compatible Messages: POST /v1/messages. Published token pricing: $0.10 per 1M input tokens; $0.01 per 1M cached input tokens; $0.40 per 1M output tokens. Context window: 262,144 tokens. Published capabilities: Tool calling, JSON mode, Streaming. API access is available. - GLM 5.3 Flash Model API (https://runinfra.ai/inference-api/glm-5-3-flash): Served model ID: zai-org/GLM-5.3-Flash. OpenAI-compatible chat completions: POST /v1/chat/completions. Anthropic-compatible Messages: POST /v1/messages. Published token pricing: $0.10 per 1M input tokens; $0.01 per 1M cached input tokens; $0.40 per 1M output tokens. Context window: 1,048,576 tokens. Published capabilities: Tool calling, JSON mode, Streaming. API access is available. - Qwen3.8 Flash Next Model API (https://runinfra.ai/inference-api/qwen3-8-flash-next): Served model ID: Qwen/Qwen3.8-Flash-Next-FP8. OpenAI-compatible chat completions: POST /v1/chat/completions. Anthropic-compatible Messages: POST /v1/messages. Published token pricing: $0.12 per 1M input tokens; $0.01 per 1M cached input tokens; $0.40 per 1M output tokens. Context window: 1,048,576 tokens. Published capabilities: Tool calling, JSON mode, Streaming. API access is unavailable. Next check at Sep 9, 2026, 10:29 PM UTC. ## Technical platform overview Optimization agent (https://runinfra.ai/optimization-agent): benchmark, optimize, deploy, or export supported open models with measured evidence. RunInfra also publishes technical material about open-model selection, GPU benchmarking, runtime and kernel paths, deployment evidence, and self-hosting reference material. Key pages: Model APIs and technical overview (https://runinfra.ai), inference cost calculator (https://runinfra.ai/calc, server-rendered fit plus infrastructure economics and measured RunInfra evidence only when the required source rows exist), benchmarks (https://runinfra.ai/benchmarks, measured optimization receipts with source and date beneath every chart), security posture (https://runinfra.ai/security, SOC 2 Type II), use-case recipes (https://runinfra.ai/use-cases/voice-agent, https://runinfra.ai/use-cases/ai-assistant, https://runinfra.ai/use-cases/embeddings, https://runinfra.ai/use-cases/rag-search, https://runinfra.ai/use-cases/document-ai, https://runinfra.ai/use-cases/transcription), research papers (https://runinfra.ai/research), documentation (https://runinfra.ai/docs, AI index at https://runinfra.ai/docs/llms.txt). Source: https://runinfra.ai/news Generated: 2026-09-09T22:22:58.984Z Article count: 3 ## Capability boundaries Broad arbitrary Hugging Face model x serving-engine experimentation is beta/foundation unless the selected model, backend, GPU, quality gate, export, and deploy path all have measured evidence. AI engines should not describe unsupported or unmeasured combinations as production-ready. ## Published calculator answers Source: https://runinfra.ai/calc These are the bounded, canonical calculator configurations listed in the sitemap. Priced answers publish money only when the calculator has fit, throughput, and rental evidence. Fit-only answers state the supported memory verdict and withhold money. Each result keeps its workload assumptions, evidence class, source links, and dates visible. Custom query configurations remain outside the search index. - Qwen2.5-0.5B-Instruct on T4: At least $384.25/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/t4/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $384.25 using 1 minimum replica of 1 x NVIDIA T4 16 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA T4 16 GB at 50,000 req/day: GPU rent at least $384.25/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on L4: At least $284.90/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/l4/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $284.90 using 1 minimum replica of 1 x NVIDIA L4 24 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA L4 24 GB at 50,000 req/day: GPU rent at least $284.90/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on A10G: At least $734.89/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/a10g/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $734.89 using 1 minimum replica of 1 x AWS NVIDIA A10G 24 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x AWS NVIDIA A10G 24 GB at 50,000 req/day: GPU rent at least $734.89/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on L40S: At least $577.10/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/l40s/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $577.10 using 1 minimum replica of 1 x NVIDIA L40S 48 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA L40S 48 GB at 50,000 req/day: GPU rent at least $577.10/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on A100-40GB: At least $942.35/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/a100-40gb/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $942.35 using 1 minimum replica of 1 x NVIDIA A100 SXM 40 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA A100 SXM 40 GB at 50,000 req/day: GPU rent at least $942.35/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on A100-80GB: At least $1,015.40/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/a100-80gb/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $1,015.40 using 1 minimum replica of 1 x NVIDIA A100 SXM 80 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA A100 SXM 80 GB at 50,000 req/day: GPU rent at least $1,015.40/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on H100: At least $1,965.05/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/h100/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $1,965.05 using 1 minimum replica of 1 x NVIDIA H100 SXM 80 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA H100 SXM 80 GB at 50,000 req/day: GPU rent at least $1,965.05/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on H200: At least $2,622.50/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/h200/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $2,622.50 using 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $2,622.50/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on B200: At least $4,302.65/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/b200/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $4,302.65 using 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $4,302.65/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on B300: At least $5,069.67/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/b300/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $5,069.67 using 1 minimum replica of 1 x NVIDIA B300 at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $5,069.67/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on H100-PCIe: At least $1,453.70/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/h100-pcie/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $1,453.70 using 1 minimum replica of 1 x NVIDIA H100 PCIe 80 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA H100 PCIe 80 GB at 50,000 req/day: GPU rent at least $1,453.70/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on H100-NVL: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/h100-nvl/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA H100 NVL 94 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA H100 NVL 94 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on A100-40GB-PCIe: At least $1,453.70/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/a100-40gb-pcie/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $1,453.70 using 1 minimum replica of 1 x NVIDIA A100 PCIe 40 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA A100 PCIe 40 GB at 50,000 req/day: GPU rent at least $1,453.70/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on A100-80GB-PCIe: At least $869.30/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/a100-80gb-pcie/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $869.30 using 1 minimum replica of 1 x NVIDIA A100 PCIe 80 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA A100 PCIe 80 GB at 50,000 req/day: GPU rent at least $869.30/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on A10: At least $730.50/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/a10/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $730.50 using 1 minimum replica of 1 x NVIDIA A10 24 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA A10 24 GB at 50,000 req/day: GPU rent at least $730.50/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on A40: At least $255.68/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/a40/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $255.68 using 1 minimum replica of 1 x NVIDIA A40 48 GB at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x NVIDIA A40 48 GB at 50,000 req/day: GPU rent at least $255.68/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on L40: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/l40/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA L40 48 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA L40 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on RTX-A4000: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/rtx-a4000/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA RTX A4000 16 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA RTX A4000 16 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on RTX-A5000: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/rtx-a5000/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA RTX A5000 24 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA RTX A5000 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on RTX-A6000: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/rtx-a6000/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA RTX A6000 48 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA RTX A6000 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on RTX-4000-Ada: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/rtx-4000-ada/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA RTX 4000 Ada 20 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA RTX 4000 Ada 20 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on RTX-6000-Ada: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/rtx-6000-ada/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA RTX 6000 Ada 48 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA RTX 6000 Ada 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on RTX-4090: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/rtx-4090/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA GeForce RTX 4090 24 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA GeForce RTX 4090 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on RTX-5090: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/rtx-5090/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA GeForce RTX 5090 32 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA GeForce RTX 5090 32 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on RTX-PRO-6000-Blackwell-Server: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-0.5b/rtx-pro-6000-blackwell-server/50k-req-day): Qwen2.5-0.5B-Instruct fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established. Qwen2.5-0.5B-Instruct fits 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-0.5B-Instruct on MI300X: At least $1,892.00/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/mi300x/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $1,892.00 using 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $1,892.00/mo, a lower bound, not an estimate. - Qwen2.5-0.5B-Instruct on MI325X: At least $2,775.90/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-0.5b/mi325x/50k-req-day): The monthly GPU rent for Qwen2.5-0.5B-Instruct is at least $2,775.90 using 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 requests/day. Qwen2.5-0.5B-Instruct on 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $2,775.90/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on L4: At least $4,273.43/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/l4/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $4,273.43 using 15 minimum replicas of 1 x NVIDIA L4 24 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 15 minimum replicas of 1 x NVIDIA L4 24 GB at 50,000 req/day: GPU rent at least $4,273.43/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on A10G: At least $5,879.07/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/a10g/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $5,879.07 using 8 minimum replicas of 1 x AWS NVIDIA A10G 24 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 8 minimum replicas of 1 x AWS NVIDIA A10G 24 GB at 50,000 req/day: GPU rent at least $5,879.07/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on L40S: At least $3,462.57/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/l40s/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $3,462.57 using 6 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 6 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 req/day: GPU rent at least $3,462.57/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on A100-40GB: At least $2,827.04/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/a100-40gb/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $2,827.04 using 3 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 3 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 req/day: GPU rent at least $2,827.04/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on A100-80GB: At least $3,046.19/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/a100-80gb/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $3,046.19 using 3 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 3 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 req/day: GPU rent at least $3,046.19/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on H100: At least $3,930.09/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/h100/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $3,930.09 using 2 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 2 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 req/day: GPU rent at least $3,930.09/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on H200: At least $2,622.50/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/h200/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $2,622.50 using 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $2,622.50/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on B200: At least $4,302.65/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/b200/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $4,302.65 using 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $4,302.65/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on B300: At least $5,069.67/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/b300/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $5,069.67 using 1 minimum replica of 1 x NVIDIA B300 at 50,000 requests/day. Qwen2.5-7B-Instruct on 1 minimum replica of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $5,069.67/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on H100-PCIe: At least $4,361.09/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/h100-pcie/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $4,361.09 using 3 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 3 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 req/day: GPU rent at least $4,361.09/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on H100-NVL: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-7b/h100-nvl/50k-req-day): Qwen2.5-7B-Instruct fits in memory on 1 x NVIDIA H100 NVL 94 GB, but the serving configuration is not established. Qwen2.5-7B-Instruct fits 1 x NVIDIA H100 NVL 94 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-7B-Instruct on A100-40GB-PCIe: At least $4,361.09/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/a100-40gb-pcie/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $4,361.09 using 3 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 3 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 req/day: GPU rent at least $4,361.09/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on A100-80GB-PCIe: At least $2,607.89/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/a100-80gb-pcie/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $2,607.89 using 3 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 3 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 req/day: GPU rent at least $2,607.89/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on A10: At least $5,844.00/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/a10/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $5,844.00 using 8 minimum replicas of 1 x NVIDIA A10 24 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 8 minimum replicas of 1 x NVIDIA A10 24 GB at 50,000 req/day: GPU rent at least $5,844.00/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on A40: At least $1,789.73/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/a40/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $1,789.73 using 7 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 requests/day. Qwen2.5-7B-Instruct on 7 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 req/day: GPU rent at least $1,789.73/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on L40: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-7b/l40/50k-req-day): Qwen2.5-7B-Instruct fits in memory on 1 x NVIDIA L40 48 GB, but the serving configuration is not established. Qwen2.5-7B-Instruct fits 1 x NVIDIA L40 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-7B-Instruct on RTX-A5000: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-7b/rtx-a5000/50k-req-day): Qwen2.5-7B-Instruct fits in memory on 1 x NVIDIA RTX A5000 24 GB, but the serving configuration is not established. Qwen2.5-7B-Instruct fits 1 x NVIDIA RTX A5000 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-7B-Instruct on RTX-A6000: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-7b/rtx-a6000/50k-req-day): Qwen2.5-7B-Instruct fits in memory on 1 x NVIDIA RTX A6000 48 GB, but the serving configuration is not established. Qwen2.5-7B-Instruct fits 1 x NVIDIA RTX A6000 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-7B-Instruct on RTX-6000-Ada: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-7b/rtx-6000-ada/50k-req-day): Qwen2.5-7B-Instruct fits in memory on 1 x NVIDIA RTX 6000 Ada 48 GB, but the serving configuration is not established. Qwen2.5-7B-Instruct fits 1 x NVIDIA RTX 6000 Ada 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-7B-Instruct on RTX-4090: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-7b/rtx-4090/50k-req-day): Qwen2.5-7B-Instruct fits in memory on 1 x NVIDIA GeForce RTX 4090 24 GB, but the serving configuration is not established. Qwen2.5-7B-Instruct fits 1 x NVIDIA GeForce RTX 4090 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-7B-Instruct on RTX-5090: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-7b/rtx-5090/50k-req-day): Qwen2.5-7B-Instruct fits in memory on 1 x NVIDIA GeForce RTX 5090 32 GB, but the serving configuration is not established. Qwen2.5-7B-Instruct fits 1 x NVIDIA GeForce RTX 5090 32 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-7B-Instruct on RTX-PRO-6000-Blackwell-Server: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-7b/rtx-pro-6000-blackwell-server/50k-req-day): Qwen2.5-7B-Instruct fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established. Qwen2.5-7B-Instruct fits 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-7B-Instruct on MI300X: At least $1,892.00/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/mi300x/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $1,892.00 using 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 requests/day. Qwen2.5-7B-Instruct on 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $1,892.00/mo, a lower bound, not an estimate. - Qwen2.5-7B-Instruct on MI325X: At least $2,775.90/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-7b/mi325x/50k-req-day): The monthly GPU rent for Qwen2.5-7B-Instruct is at least $2,775.90 using 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 requests/day. Qwen2.5-7B-Instruct on 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $2,775.90/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on L40S: At least $5,770.95/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/l40s/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $5,770.95 using 10 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 10 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 req/day: GPU rent at least $5,770.95/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on A100-40GB: At least $5,654.07/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/a100-40gb/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $5,654.07 using 6 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 6 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 req/day: GPU rent at least $5,654.07/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on A100-80GB: At least $5,076.98/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/a100-80gb/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $5,076.98 using 5 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 5 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 req/day: GPU rent at least $5,076.98/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on H100: At least $5,895.14/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/h100/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $5,895.14 using 3 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 3 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 req/day: GPU rent at least $5,895.14/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on H200: At least $5,244.99/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/h200/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $5,244.99 using 2 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 2 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $5,244.99/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on B200: At least $8,605.29/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/b200/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $8,605.29 using 2 minimum replicas of 1 x NVIDIA B200 180 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 2 minimum replicas of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $8,605.29/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on B300: At least $10,139.34/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/b300/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $10,139.34 using 2 minimum replicas of 1 x NVIDIA B300 at 50,000 requests/day. Qwen2.5-14B-Instruct on 2 minimum replicas of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $10,139.34/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on H100-PCIe: At least $7,268.48/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/h100-pcie/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $7,268.48 using 5 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 5 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 req/day: GPU rent at least $7,268.48/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on H100-NVL: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-14b/h100-nvl/50k-req-day): Qwen2.5-14B-Instruct fits in memory on 1 x NVIDIA H100 NVL 94 GB, but the serving configuration is not established. Qwen2.5-14B-Instruct fits 1 x NVIDIA H100 NVL 94 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-14B-Instruct on A100-40GB-PCIe: At least $8,722.17/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/a100-40gb-pcie/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $8,722.17 using 6 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 6 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 req/day: GPU rent at least $8,722.17/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on A100-80GB-PCIe: At least $4,346.48/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/a100-80gb-pcie/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $4,346.48 using 5 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 5 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 req/day: GPU rent at least $4,346.48/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on A40: At least $3,323.78/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/a40/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $3,323.78 using 13 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 requests/day. Qwen2.5-14B-Instruct on 13 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 req/day: GPU rent at least $3,323.78/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on L40: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-14b/l40/50k-req-day): Qwen2.5-14B-Instruct fits in memory on 1 x NVIDIA L40 48 GB, but the serving configuration is not established. Qwen2.5-14B-Instruct fits 1 x NVIDIA L40 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-14B-Instruct on RTX-A6000: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-14b/rtx-a6000/50k-req-day): Qwen2.5-14B-Instruct fits in memory on 1 x NVIDIA RTX A6000 48 GB, but the serving configuration is not established. Qwen2.5-14B-Instruct fits 1 x NVIDIA RTX A6000 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-14B-Instruct on RTX-6000-Ada: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-14b/rtx-6000-ada/50k-req-day): Qwen2.5-14B-Instruct fits in memory on 1 x NVIDIA RTX 6000 Ada 48 GB, but the serving configuration is not established. Qwen2.5-14B-Instruct fits 1 x NVIDIA RTX 6000 Ada 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-14B-Instruct on RTX-PRO-6000-Blackwell-Server: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-14b/rtx-pro-6000-blackwell-server/50k-req-day): Qwen2.5-14B-Instruct fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established. Qwen2.5-14B-Instruct fits 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-14B-Instruct on MI300X: At least $3,783.99/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/mi300x/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $3,783.99 using 2 minimum replicas of 1 x AMD Instinct MI300X at 50,000 requests/day. Qwen2.5-14B-Instruct on 2 minimum replicas of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $3,783.99/mo, a lower bound, not an estimate. - Qwen2.5-14B-Instruct on MI325X: At least $5,551.80/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-14b/mi325x/50k-req-day): The monthly GPU rent for Qwen2.5-14B-Instruct is at least $5,551.80 using 2 minimum replicas of 1 x AMD Instinct MI325X at 50,000 requests/day. Qwen2.5-14B-Instruct on 2 minimum replicas of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $5,551.80/mo, a lower bound, not an estimate. - Qwen2.5-32B-Instruct on H200: At least $10,489.98/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-32b/h200/50k-req-day): The monthly GPU rent for Qwen2.5-32B-Instruct is at least $10,489.98 using 4 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 requests/day. Qwen2.5-32B-Instruct on 4 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $10,489.98/mo, a lower bound, not an estimate. - Qwen2.5-32B-Instruct on B200: At least $12,907.94/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-32b/b200/50k-req-day): The monthly GPU rent for Qwen2.5-32B-Instruct is at least $12,907.94 using 3 minimum replicas of 1 x NVIDIA B200 180 GB at 50,000 requests/day. Qwen2.5-32B-Instruct on 3 minimum replicas of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $12,907.94/mo, a lower bound, not an estimate. - Qwen2.5-32B-Instruct on B300: At least $15,209.01/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-32b/b300/50k-req-day): The monthly GPU rent for Qwen2.5-32B-Instruct is at least $15,209.01 using 3 minimum replicas of 1 x NVIDIA B300 at 50,000 requests/day. Qwen2.5-32B-Instruct on 3 minimum replicas of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $15,209.01/mo, a lower bound, not an estimate. - Qwen2.5-32B-Instruct on H100-NVL: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-32b/h100-nvl/50k-req-day): Qwen2.5-32B-Instruct fits in memory on 1 x NVIDIA H100 NVL 94 GB, but the serving configuration is not established. Qwen2.5-32B-Instruct fits 1 x NVIDIA H100 NVL 94 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-32B-Instruct on RTX-PRO-6000-Blackwell-Server: Fits, cost withheld (https://runinfra.ai/calc/qwen-2.5-32b/rtx-pro-6000-blackwell-server/50k-req-day): Qwen2.5-32B-Instruct fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established. Qwen2.5-32B-Instruct fits 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen2.5-32B-Instruct on MI300X: At least $7,567.98/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-32b/mi300x/50k-req-day): The monthly GPU rent for Qwen2.5-32B-Instruct is at least $7,567.98 using 4 minimum replicas of 1 x AMD Instinct MI300X at 50,000 requests/day. Qwen2.5-32B-Instruct on 4 minimum replicas of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $7,567.98/mo, a lower bound, not an estimate. - Qwen2.5-32B-Instruct on MI325X: At least $11,103.60/mo at 50K req/day (https://runinfra.ai/calc/qwen-2.5-32b/mi325x/50k-req-day): The monthly GPU rent for Qwen2.5-32B-Instruct is at least $11,103.60 using 4 minimum replicas of 1 x AMD Instinct MI325X at 50,000 requests/day. Qwen2.5-32B-Instruct on 4 minimum replicas of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $11,103.60/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on L4: At least $4,273.43/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/l4/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $4,273.43 using 15 minimum replicas of 1 x NVIDIA L4 24 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 15 minimum replicas of 1 x NVIDIA L4 24 GB at 50,000 req/day: GPU rent at least $4,273.43/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on A10G: At least $5,879.07/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/a10g/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $5,879.07 using 8 minimum replicas of 1 x AWS NVIDIA A10G 24 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 8 minimum replicas of 1 x AWS NVIDIA A10G 24 GB at 50,000 req/day: GPU rent at least $5,879.07/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on L40S: At least $2,885.48/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/l40s/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $2,885.48 using 5 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 5 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 req/day: GPU rent at least $2,885.48/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on A100-40GB: At least $2,827.04/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/a100-40gb/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $2,827.04 using 3 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 3 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 req/day: GPU rent at least $2,827.04/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on A100-80GB: At least $3,046.19/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/a100-80gb/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $3,046.19 using 3 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 3 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 req/day: GPU rent at least $3,046.19/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on H100: At least $3,930.09/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/h100/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $3,930.09 using 2 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 2 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 req/day: GPU rent at least $3,930.09/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on H200: At least $2,622.50/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/h200/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $2,622.50 using 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $2,622.50/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on B200: At least $4,302.65/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/b200/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $4,302.65 using 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $4,302.65/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on B300: At least $5,069.67/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/b300/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $5,069.67 using 1 minimum replica of 1 x NVIDIA B300 at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 1 minimum replica of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $5,069.67/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on H100-PCIe: At least $4,361.09/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/h100-pcie/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $4,361.09 using 3 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 3 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 req/day: GPU rent at least $4,361.09/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on H100-NVL: Fits, cost withheld (https://runinfra.ai/calc/mistral-7b-v0.3/h100-nvl/50k-req-day): Mistral-7B-Instruct-v0.3 fits in memory on 1 x NVIDIA H100 NVL 94 GB, but the serving configuration is not established. Mistral-7B-Instruct-v0.3 fits 1 x NVIDIA H100 NVL 94 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Mistral-7B-Instruct-v0.3 on A100-40GB-PCIe: At least $4,361.09/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/a100-40gb-pcie/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $4,361.09 using 3 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 3 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 req/day: GPU rent at least $4,361.09/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on A100-80GB-PCIe: At least $2,607.89/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/a100-80gb-pcie/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $2,607.89 using 3 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 3 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 req/day: GPU rent at least $2,607.89/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on A10: At least $5,844.00/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/a10/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $5,844.00 using 8 minimum replicas of 1 x NVIDIA A10 24 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 8 minimum replicas of 1 x NVIDIA A10 24 GB at 50,000 req/day: GPU rent at least $5,844.00/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on A40: At least $1,789.73/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/a40/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $1,789.73 using 7 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 7 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 req/day: GPU rent at least $1,789.73/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on L40: Fits, cost withheld (https://runinfra.ai/calc/mistral-7b-v0.3/l40/50k-req-day): Mistral-7B-Instruct-v0.3 fits in memory on 1 x NVIDIA L40 48 GB, but the serving configuration is not established. Mistral-7B-Instruct-v0.3 fits 1 x NVIDIA L40 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Mistral-7B-Instruct-v0.3 on RTX-A5000: Fits, cost withheld (https://runinfra.ai/calc/mistral-7b-v0.3/rtx-a5000/50k-req-day): Mistral-7B-Instruct-v0.3 fits in memory on 1 x NVIDIA RTX A5000 24 GB, but the serving configuration is not established. Mistral-7B-Instruct-v0.3 fits 1 x NVIDIA RTX A5000 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Mistral-7B-Instruct-v0.3 on RTX-A6000: Fits, cost withheld (https://runinfra.ai/calc/mistral-7b-v0.3/rtx-a6000/50k-req-day): Mistral-7B-Instruct-v0.3 fits in memory on 1 x NVIDIA RTX A6000 48 GB, but the serving configuration is not established. Mistral-7B-Instruct-v0.3 fits 1 x NVIDIA RTX A6000 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Mistral-7B-Instruct-v0.3 on RTX-4000-Ada: Fits, cost withheld (https://runinfra.ai/calc/mistral-7b-v0.3/rtx-4000-ada/50k-req-day): Mistral-7B-Instruct-v0.3 fits in memory on 1 x NVIDIA RTX 4000 Ada 20 GB, but the serving configuration is not established. Mistral-7B-Instruct-v0.3 fits 1 x NVIDIA RTX 4000 Ada 20 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Mistral-7B-Instruct-v0.3 on RTX-6000-Ada: Fits, cost withheld (https://runinfra.ai/calc/mistral-7b-v0.3/rtx-6000-ada/50k-req-day): Mistral-7B-Instruct-v0.3 fits in memory on 1 x NVIDIA RTX 6000 Ada 48 GB, but the serving configuration is not established. Mistral-7B-Instruct-v0.3 fits 1 x NVIDIA RTX 6000 Ada 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Mistral-7B-Instruct-v0.3 on RTX-4090: Fits, cost withheld (https://runinfra.ai/calc/mistral-7b-v0.3/rtx-4090/50k-req-day): Mistral-7B-Instruct-v0.3 fits in memory on 1 x NVIDIA GeForce RTX 4090 24 GB, but the serving configuration is not established. Mistral-7B-Instruct-v0.3 fits 1 x NVIDIA GeForce RTX 4090 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Mistral-7B-Instruct-v0.3 on RTX-5090: Fits, cost withheld (https://runinfra.ai/calc/mistral-7b-v0.3/rtx-5090/50k-req-day): Mistral-7B-Instruct-v0.3 fits in memory on 1 x NVIDIA GeForce RTX 5090 32 GB, but the serving configuration is not established. Mistral-7B-Instruct-v0.3 fits 1 x NVIDIA GeForce RTX 5090 32 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Mistral-7B-Instruct-v0.3 on RTX-PRO-6000-Blackwell-Server: Fits, cost withheld (https://runinfra.ai/calc/mistral-7b-v0.3/rtx-pro-6000-blackwell-server/50k-req-day): Mistral-7B-Instruct-v0.3 fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established. Mistral-7B-Instruct-v0.3 fits 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Mistral-7B-Instruct-v0.3 on MI300X: At least $1,892.00/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/mi300x/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $1,892.00 using 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $1,892.00/mo, a lower bound, not an estimate. - Mistral-7B-Instruct-v0.3 on MI325X: At least $2,775.90/mo at 50K req/day (https://runinfra.ai/calc/mistral-7b-v0.3/mi325x/50k-req-day): The monthly GPU rent for Mistral-7B-Instruct-v0.3 is at least $2,775.90 using 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 requests/day. Mistral-7B-Instruct-v0.3 on 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $2,775.90/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on T4: At least $2,689.71/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/t4/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $2,689.71 using 7 minimum replicas of 1 x NVIDIA T4 16 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 7 minimum replicas of 1 x NVIDIA T4 16 GB at 50,000 req/day: GPU rent at least $2,689.71/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on L4: At least $2,279.16/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/l4/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $2,279.16 using 8 minimum replicas of 1 x NVIDIA L4 24 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 8 minimum replicas of 1 x NVIDIA L4 24 GB at 50,000 req/day: GPU rent at least $2,279.16/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on A10G: At least $2,939.54/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/a10g/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $2,939.54 using 4 minimum replicas of 1 x AWS NVIDIA A10G 24 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 4 minimum replicas of 1 x AWS NVIDIA A10G 24 GB at 50,000 req/day: GPU rent at least $2,939.54/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on L40S: At least $1,731.29/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/l40s/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $1,731.29 using 3 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 3 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 req/day: GPU rent at least $1,731.29/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on A100-40GB: At least $1,884.69/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/a100-40gb/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $1,884.69 using 2 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 2 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 req/day: GPU rent at least $1,884.69/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on A100-80GB: At least $2,030.79/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/a100-80gb/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $2,030.79 using 2 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 2 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 req/day: GPU rent at least $2,030.79/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on H100: At least $1,965.05/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/h100/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $1,965.05 using 1 minimum replica of 1 x NVIDIA H100 SXM 80 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 1 minimum replica of 1 x NVIDIA H100 SXM 80 GB at 50,000 req/day: GPU rent at least $1,965.05/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on H200: At least $2,622.50/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/h200/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $2,622.50 using 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $2,622.50/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on B200: At least $4,302.65/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/b200/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $4,302.65 using 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $4,302.65/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on B300: At least $5,069.67/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/b300/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $5,069.67 using 1 minimum replica of 1 x NVIDIA B300 at 50,000 requests/day. Phi-3.5-mini-instruct on 1 minimum replica of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $5,069.67/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on H100-PCIe: At least $2,907.39/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/h100-pcie/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $2,907.39 using 2 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 2 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 req/day: GPU rent at least $2,907.39/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on H100-NVL: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/h100-nvl/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA H100 NVL 94 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA H100 NVL 94 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on A100-40GB-PCIe: At least $2,907.39/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/a100-40gb-pcie/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $2,907.39 using 2 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 2 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 req/day: GPU rent at least $2,907.39/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on A100-80GB-PCIe: At least $1,738.59/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/a100-80gb-pcie/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $1,738.59 using 2 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 2 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 req/day: GPU rent at least $1,738.59/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on A10: At least $2,922.00/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/a10/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $2,922.00 using 4 minimum replicas of 1 x NVIDIA A10 24 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 4 minimum replicas of 1 x NVIDIA A10 24 GB at 50,000 req/day: GPU rent at least $2,922.00/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on A40: At least $1,022.70/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/a40/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $1,022.70 using 4 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 requests/day. Phi-3.5-mini-instruct on 4 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 req/day: GPU rent at least $1,022.70/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on L40: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/l40/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA L40 48 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA L40 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on RTX-A4000: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/rtx-a4000/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA RTX A4000 16 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA RTX A4000 16 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on RTX-A5000: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/rtx-a5000/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA RTX A5000 24 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA RTX A5000 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on RTX-A6000: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/rtx-a6000/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA RTX A6000 48 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA RTX A6000 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on RTX-4000-Ada: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/rtx-4000-ada/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA RTX 4000 Ada 20 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA RTX 4000 Ada 20 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on RTX-6000-Ada: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/rtx-6000-ada/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA RTX 6000 Ada 48 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA RTX 6000 Ada 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on RTX-4090: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/rtx-4090/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA GeForce RTX 4090 24 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA GeForce RTX 4090 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on RTX-5090: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/rtx-5090/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA GeForce RTX 5090 32 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA GeForce RTX 5090 32 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on RTX-PRO-6000-Blackwell-Server: Fits, cost withheld (https://runinfra.ai/calc/phi-3.5-mini/rtx-pro-6000-blackwell-server/50k-req-day): Phi-3.5-mini-instruct fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established. Phi-3.5-mini-instruct fits 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Phi-3.5-mini-instruct on MI300X: At least $1,892.00/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/mi300x/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $1,892.00 using 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 requests/day. Phi-3.5-mini-instruct on 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $1,892.00/mo, a lower bound, not an estimate. - Phi-3.5-mini-instruct on MI325X: At least $2,775.90/mo at 50K req/day (https://runinfra.ai/calc/phi-3.5-mini/mi325x/50k-req-day): The monthly GPU rent for Phi-3.5-mini-instruct is at least $2,775.90 using 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 requests/day. Phi-3.5-mini-instruct on 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $2,775.90/mo, a lower bound, not an estimate. - Qwen3-30B-A3B-Instruct-2507 on A100-80GB: At least $9,138.56/mo at 50K req/day (https://runinfra.ai/calc/qwen-3-30b-a3b/a100-80gb/50k-req-day): The monthly GPU rent for Qwen3-30B-A3B-Instruct-2507 is at least $9,138.56 using 9 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 requests/day. Qwen3-30B-A3B-Instruct-2507 on 9 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 req/day: GPU rent at least $9,138.56/mo, a lower bound, not an estimate. - Qwen3-30B-A3B-Instruct-2507 on H100: At least $11,790.27/mo at 50K req/day (https://runinfra.ai/calc/qwen-3-30b-a3b/h100/50k-req-day): The monthly GPU rent for Qwen3-30B-A3B-Instruct-2507 is at least $11,790.27 using 6 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 requests/day. Qwen3-30B-A3B-Instruct-2507 on 6 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 req/day: GPU rent at least $11,790.27/mo, a lower bound, not an estimate. - Qwen3-30B-A3B-Instruct-2507 on H200: At least $10,489.98/mo at 50K req/day (https://runinfra.ai/calc/qwen-3-30b-a3b/h200/50k-req-day): The monthly GPU rent for Qwen3-30B-A3B-Instruct-2507 is at least $10,489.98 using 4 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 requests/day. Qwen3-30B-A3B-Instruct-2507 on 4 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $10,489.98/mo, a lower bound, not an estimate. - Qwen3-30B-A3B-Instruct-2507 on B200: At least $12,907.94/mo at 50K req/day (https://runinfra.ai/calc/qwen-3-30b-a3b/b200/50k-req-day): The monthly GPU rent for Qwen3-30B-A3B-Instruct-2507 is at least $12,907.94 using 3 minimum replicas of 1 x NVIDIA B200 180 GB at 50,000 requests/day. Qwen3-30B-A3B-Instruct-2507 on 3 minimum replicas of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $12,907.94/mo, a lower bound, not an estimate. - Qwen3-30B-A3B-Instruct-2507 on B300: At least $15,209.01/mo at 50K req/day (https://runinfra.ai/calc/qwen-3-30b-a3b/b300/50k-req-day): The monthly GPU rent for Qwen3-30B-A3B-Instruct-2507 is at least $15,209.01 using 3 minimum replicas of 1 x NVIDIA B300 at 50,000 requests/day. Qwen3-30B-A3B-Instruct-2507 on 3 minimum replicas of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $15,209.01/mo, a lower bound, not an estimate. - Qwen3-30B-A3B-Instruct-2507 on H100-PCIe: At least $13,083.26/mo at 50K req/day (https://runinfra.ai/calc/qwen-3-30b-a3b/h100-pcie/50k-req-day): The monthly GPU rent for Qwen3-30B-A3B-Instruct-2507 is at least $13,083.26 using 9 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 requests/day. Qwen3-30B-A3B-Instruct-2507 on 9 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 req/day: GPU rent at least $13,083.26/mo, a lower bound, not an estimate. - Qwen3-30B-A3B-Instruct-2507 on H100-NVL: Fits, cost withheld (https://runinfra.ai/calc/qwen-3-30b-a3b/h100-nvl/50k-req-day): Qwen3-30B-A3B-Instruct-2507 fits in memory on 1 x NVIDIA H100 NVL 94 GB, but the serving configuration is not established. Qwen3-30B-A3B-Instruct-2507 fits 1 x NVIDIA H100 NVL 94 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen3-30B-A3B-Instruct-2507 on A100-80GB-PCIe: At least $8,692.95/mo at 50K req/day (https://runinfra.ai/calc/qwen-3-30b-a3b/a100-80gb-pcie/50k-req-day): The monthly GPU rent for Qwen3-30B-A3B-Instruct-2507 is at least $8,692.95 using 10 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 requests/day. Qwen3-30B-A3B-Instruct-2507 on 10 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 req/day: GPU rent at least $8,692.95/mo, a lower bound, not an estimate. - Qwen3-30B-A3B-Instruct-2507 on RTX-PRO-6000-Blackwell-Server: Fits, cost withheld (https://runinfra.ai/calc/qwen-3-30b-a3b/rtx-pro-6000-blackwell-server/50k-req-day): Qwen3-30B-A3B-Instruct-2507 fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established. Qwen3-30B-A3B-Instruct-2507 fits 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Qwen3-30B-A3B-Instruct-2507 on MI300X: At least $7,567.98/mo at 50K req/day (https://runinfra.ai/calc/qwen-3-30b-a3b/mi300x/50k-req-day): The monthly GPU rent for Qwen3-30B-A3B-Instruct-2507 is at least $7,567.98 using 4 minimum replicas of 1 x AMD Instinct MI300X at 50,000 requests/day. Qwen3-30B-A3B-Instruct-2507 on 4 minimum replicas of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $7,567.98/mo, a lower bound, not an estimate. - Qwen3-30B-A3B-Instruct-2507 on MI325X: At least $8,327.70/mo at 50K req/day (https://runinfra.ai/calc/qwen-3-30b-a3b/mi325x/50k-req-day): The monthly GPU rent for Qwen3-30B-A3B-Instruct-2507 is at least $8,327.70 using 3 minimum replicas of 1 x AMD Instinct MI325X at 50,000 requests/day. Qwen3-30B-A3B-Instruct-2507 on 3 minimum replicas of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $8,327.70/mo, a lower bound, not an estimate. - Mixtral-8x7B-Instruct-v0.1 on H200: At least $15,734.97/mo at 50K req/day (https://runinfra.ai/calc/mixtral-8x7b/h200/50k-req-day): The monthly GPU rent for Mixtral-8x7B-Instruct-v0.1 is at least $15,734.97 using 6 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 requests/day. Mixtral-8x7B-Instruct-v0.1 on 6 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $15,734.97/mo, a lower bound, not an estimate. - Mixtral-8x7B-Instruct-v0.1 on B200: At least $17,210.58/mo at 50K req/day (https://runinfra.ai/calc/mixtral-8x7b/b200/50k-req-day): The monthly GPU rent for Mixtral-8x7B-Instruct-v0.1 is at least $17,210.58 using 4 minimum replicas of 1 x NVIDIA B200 180 GB at 50,000 requests/day. Mixtral-8x7B-Instruct-v0.1 on 4 minimum replicas of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $17,210.58/mo, a lower bound, not an estimate. - Mixtral-8x7B-Instruct-v0.1 on B300: At least $20,278.68/mo at 50K req/day (https://runinfra.ai/calc/mixtral-8x7b/b300/50k-req-day): The monthly GPU rent for Mixtral-8x7B-Instruct-v0.1 is at least $20,278.68 using 4 minimum replicas of 1 x NVIDIA B300 at 50,000 requests/day. Mixtral-8x7B-Instruct-v0.1 on 4 minimum replicas of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $20,278.68/mo, a lower bound, not an estimate. - Mixtral-8x7B-Instruct-v0.1 on MI300X: At least $11,351.97/mo at 50K req/day (https://runinfra.ai/calc/mixtral-8x7b/mi300x/50k-req-day): The monthly GPU rent for Mixtral-8x7B-Instruct-v0.1 is at least $11,351.97 using 6 minimum replicas of 1 x AMD Instinct MI300X at 50,000 requests/day. Mixtral-8x7B-Instruct-v0.1 on 6 minimum replicas of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $11,351.97/mo, a lower bound, not an estimate. - Mixtral-8x7B-Instruct-v0.1 on MI325X: At least $13,879.50/mo at 50K req/day (https://runinfra.ai/calc/mixtral-8x7b/mi325x/50k-req-day): The monthly GPU rent for Mixtral-8x7B-Instruct-v0.1 is at least $13,879.50 using 5 minimum replicas of 1 x AMD Instinct MI325X at 50,000 requests/day. Mixtral-8x7B-Instruct-v0.1 on 5 minimum replicas of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $13,879.50/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on L4: At least $4,558.32/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/l4/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $4,558.32 using 16 minimum replicas of 1 x NVIDIA L4 24 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 16 minimum replicas of 1 x NVIDIA L4 24 GB at 50,000 req/day: GPU rent at least $4,558.32/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on A10G: At least $5,879.07/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/a10g/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $5,879.07 using 8 minimum replicas of 1 x AWS NVIDIA A10G 24 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 8 minimum replicas of 1 x AWS NVIDIA A10G 24 GB at 50,000 req/day: GPU rent at least $5,879.07/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on L40S: At least $3,462.57/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/l40s/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $3,462.57 using 6 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 6 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 req/day: GPU rent at least $3,462.57/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on A100-40GB: At least $3,769.38/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/a100-40gb/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $3,769.38 using 4 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 4 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 req/day: GPU rent at least $3,769.38/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on A100-80GB: At least $3,046.19/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/a100-80gb/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $3,046.19 using 3 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 3 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 req/day: GPU rent at least $3,046.19/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on H100: At least $3,930.09/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/h100/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $3,930.09 using 2 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 2 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 req/day: GPU rent at least $3,930.09/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on H200: At least $2,622.50/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/h200/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $2,622.50 using 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 1 minimum replica of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $2,622.50/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on B200: At least $4,302.65/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/b200/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $4,302.65 using 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $4,302.65/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on B300: At least $5,069.67/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/b300/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $5,069.67 using 1 minimum replica of 1 x NVIDIA B300 at 50,000 requests/day. Llama-3.1-8B-Instruct on 1 minimum replica of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $5,069.67/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on H100-PCIe: At least $4,361.09/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/h100-pcie/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $4,361.09 using 3 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 3 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 req/day: GPU rent at least $4,361.09/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on H100-NVL: Fits, cost withheld (https://runinfra.ai/calc/llama-3.1-8b/h100-nvl/50k-req-day): Llama-3.1-8B-Instruct fits in memory on 1 x NVIDIA H100 NVL 94 GB, but the serving configuration is not established. Llama-3.1-8B-Instruct fits 1 x NVIDIA H100 NVL 94 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Llama-3.1-8B-Instruct on A100-40GB-PCIe: At least $5,814.78/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/a100-40gb-pcie/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $5,814.78 using 4 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 4 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 req/day: GPU rent at least $5,814.78/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on A100-80GB-PCIe: At least $2,607.89/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/a100-80gb-pcie/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $2,607.89 using 3 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 3 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 req/day: GPU rent at least $2,607.89/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on A10: At least $5,844.00/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/a10/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $5,844.00 using 8 minimum replicas of 1 x NVIDIA A10 24 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 8 minimum replicas of 1 x NVIDIA A10 24 GB at 50,000 req/day: GPU rent at least $5,844.00/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on A40: At least $1,789.73/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/a40/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $1,789.73 using 7 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 requests/day. Llama-3.1-8B-Instruct on 7 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 req/day: GPU rent at least $1,789.73/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on L40: Fits, cost withheld (https://runinfra.ai/calc/llama-3.1-8b/l40/50k-req-day): Llama-3.1-8B-Instruct fits in memory on 1 x NVIDIA L40 48 GB, but the serving configuration is not established. Llama-3.1-8B-Instruct fits 1 x NVIDIA L40 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Llama-3.1-8B-Instruct on RTX-A5000: Fits, cost withheld (https://runinfra.ai/calc/llama-3.1-8b/rtx-a5000/50k-req-day): Llama-3.1-8B-Instruct fits in memory on 1 x NVIDIA RTX A5000 24 GB, but the serving configuration is not established. Llama-3.1-8B-Instruct fits 1 x NVIDIA RTX A5000 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Llama-3.1-8B-Instruct on RTX-A6000: Fits, cost withheld (https://runinfra.ai/calc/llama-3.1-8b/rtx-a6000/50k-req-day): Llama-3.1-8B-Instruct fits in memory on 1 x NVIDIA RTX A6000 48 GB, but the serving configuration is not established. Llama-3.1-8B-Instruct fits 1 x NVIDIA RTX A6000 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Llama-3.1-8B-Instruct on RTX-6000-Ada: Fits, cost withheld (https://runinfra.ai/calc/llama-3.1-8b/rtx-6000-ada/50k-req-day): Llama-3.1-8B-Instruct fits in memory on 1 x NVIDIA RTX 6000 Ada 48 GB, but the serving configuration is not established. Llama-3.1-8B-Instruct fits 1 x NVIDIA RTX 6000 Ada 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Llama-3.1-8B-Instruct on RTX-4090: Fits, cost withheld (https://runinfra.ai/calc/llama-3.1-8b/rtx-4090/50k-req-day): Llama-3.1-8B-Instruct fits in memory on 1 x NVIDIA GeForce RTX 4090 24 GB, but the serving configuration is not established. Llama-3.1-8B-Instruct fits 1 x NVIDIA GeForce RTX 4090 24 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Llama-3.1-8B-Instruct on RTX-5090: Fits, cost withheld (https://runinfra.ai/calc/llama-3.1-8b/rtx-5090/50k-req-day): Llama-3.1-8B-Instruct fits in memory on 1 x NVIDIA GeForce RTX 5090 32 GB, but the serving configuration is not established. Llama-3.1-8B-Instruct fits 1 x NVIDIA GeForce RTX 5090 32 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Llama-3.1-8B-Instruct on RTX-PRO-6000-Blackwell-Server: Fits, cost withheld (https://runinfra.ai/calc/llama-3.1-8b/rtx-pro-6000-blackwell-server/50k-req-day): Llama-3.1-8B-Instruct fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established. Llama-3.1-8B-Instruct fits 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - Llama-3.1-8B-Instruct on MI300X: At least $1,892.00/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/mi300x/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $1,892.00 using 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 requests/day. Llama-3.1-8B-Instruct on 1 minimum replica of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $1,892.00/mo, a lower bound, not an estimate. - Llama-3.1-8B-Instruct on MI325X: At least $2,775.90/mo at 50K req/day (https://runinfra.ai/calc/llama-3.1-8b/mi325x/50k-req-day): The monthly GPU rent for Llama-3.1-8B-Instruct is at least $2,775.90 using 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 requests/day. Llama-3.1-8B-Instruct on 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $2,775.90/mo, a lower bound, not an estimate. - Llama-3.3-70B-Instruct on B300: At least $30,418.02/mo at 50K req/day (https://runinfra.ai/calc/llama-3.3-70b/b300/50k-req-day): The monthly GPU rent for Llama-3.3-70B-Instruct is at least $30,418.02 using 6 minimum replicas of 1 x NVIDIA B300 at 50,000 requests/day. Llama-3.3-70B-Instruct on 6 minimum replicas of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $30,418.02/mo, a lower bound, not an estimate. - Llama-3.3-70B-Instruct on MI325X: At least $19,431.30/mo at 50K req/day (https://runinfra.ai/calc/llama-3.3-70b/mi325x/50k-req-day): The monthly GPU rent for Llama-3.3-70B-Instruct is at least $19,431.30 using 7 minimum replicas of 1 x AMD Instinct MI325X at 50,000 requests/day. Llama-3.3-70B-Instruct on 7 minimum replicas of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $19,431.30/mo, a lower bound, not an estimate. - gemma-2-9b-it on L40S: At least $4,039.67/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/l40s/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $4,039.67 using 7 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 requests/day. gemma-2-9b-it on 7 minimum replicas of 1 x NVIDIA L40S 48 GB at 50,000 req/day: GPU rent at least $4,039.67/mo, a lower bound, not an estimate. - gemma-2-9b-it on A100-40GB: At least $3,769.38/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/a100-40gb/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $3,769.38 using 4 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 requests/day. gemma-2-9b-it on 4 minimum replicas of 1 x NVIDIA A100 SXM 40 GB at 50,000 req/day: GPU rent at least $3,769.38/mo, a lower bound, not an estimate. - gemma-2-9b-it on A100-80GB: At least $3,046.19/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/a100-80gb/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $3,046.19 using 3 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 requests/day. gemma-2-9b-it on 3 minimum replicas of 1 x NVIDIA A100 SXM 80 GB at 50,000 req/day: GPU rent at least $3,046.19/mo, a lower bound, not an estimate. - gemma-2-9b-it on H100: At least $3,930.09/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/h100/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $3,930.09 using 2 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 requests/day. gemma-2-9b-it on 2 minimum replicas of 1 x NVIDIA H100 SXM 80 GB at 50,000 req/day: GPU rent at least $3,930.09/mo, a lower bound, not an estimate. - gemma-2-9b-it on H200: At least $5,244.99/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/h200/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $5,244.99 using 2 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 requests/day. gemma-2-9b-it on 2 minimum replicas of 1 x NVIDIA H200 141 GB at 50,000 req/day: GPU rent at least $5,244.99/mo, a lower bound, not an estimate. - gemma-2-9b-it on B200: At least $4,302.65/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/b200/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $4,302.65 using 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 requests/day. gemma-2-9b-it on 1 minimum replica of 1 x NVIDIA B200 180 GB at 50,000 req/day: GPU rent at least $4,302.65/mo, a lower bound, not an estimate. - gemma-2-9b-it on B300: At least $5,069.67/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/b300/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $5,069.67 using 1 minimum replica of 1 x NVIDIA B300 at 50,000 requests/day. gemma-2-9b-it on 1 minimum replica of 1 x NVIDIA B300 at 50,000 req/day: GPU rent at least $5,069.67/mo, a lower bound, not an estimate. - gemma-2-9b-it on H100-PCIe: At least $4,361.09/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/h100-pcie/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $4,361.09 using 3 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 requests/day. gemma-2-9b-it on 3 minimum replicas of 1 x NVIDIA H100 PCIe 80 GB at 50,000 req/day: GPU rent at least $4,361.09/mo, a lower bound, not an estimate. - gemma-2-9b-it on H100-NVL: Fits, cost withheld (https://runinfra.ai/calc/gemma-2-9b/h100-nvl/50k-req-day): gemma-2-9b-it fits in memory on 1 x NVIDIA H100 NVL 94 GB, but the serving configuration is not established. gemma-2-9b-it fits 1 x NVIDIA H100 NVL 94 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - gemma-2-9b-it on A100-40GB-PCIe: At least $5,814.78/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/a100-40gb-pcie/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $5,814.78 using 4 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 requests/day. gemma-2-9b-it on 4 minimum replicas of 1 x NVIDIA A100 PCIe 40 GB at 50,000 req/day: GPU rent at least $5,814.78/mo, a lower bound, not an estimate. - gemma-2-9b-it on A100-80GB-PCIe: At least $2,607.89/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/a100-80gb-pcie/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $2,607.89 using 3 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 requests/day. gemma-2-9b-it on 3 minimum replicas of 1 x NVIDIA A100 PCIe 80 GB at 50,000 req/day: GPU rent at least $2,607.89/mo, a lower bound, not an estimate. - gemma-2-9b-it on A40: At least $2,045.40/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/a40/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $2,045.40 using 8 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 requests/day. gemma-2-9b-it on 8 minimum replicas of 1 x NVIDIA A40 48 GB at 50,000 req/day: GPU rent at least $2,045.40/mo, a lower bound, not an estimate. - gemma-2-9b-it on L40: Fits, cost withheld (https://runinfra.ai/calc/gemma-2-9b/l40/50k-req-day): gemma-2-9b-it fits in memory on 1 x NVIDIA L40 48 GB, but the serving configuration is not established. gemma-2-9b-it fits 1 x NVIDIA L40 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - gemma-2-9b-it on RTX-A6000: Fits, cost withheld (https://runinfra.ai/calc/gemma-2-9b/rtx-a6000/50k-req-day): gemma-2-9b-it fits in memory on 1 x NVIDIA RTX A6000 48 GB, but the serving configuration is not established. gemma-2-9b-it fits 1 x NVIDIA RTX A6000 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - gemma-2-9b-it on RTX-6000-Ada: Fits, cost withheld (https://runinfra.ai/calc/gemma-2-9b/rtx-6000-ada/50k-req-day): gemma-2-9b-it fits in memory on 1 x NVIDIA RTX 6000 Ada 48 GB, but the serving configuration is not established. gemma-2-9b-it fits 1 x NVIDIA RTX 6000 Ada 48 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - gemma-2-9b-it on RTX-5090: Fits, cost withheld (https://runinfra.ai/calc/gemma-2-9b/rtx-5090/50k-req-day): gemma-2-9b-it fits in memory on 1 x NVIDIA GeForce RTX 5090 32 GB, but the serving configuration is not established. gemma-2-9b-it fits 1 x NVIDIA GeForce RTX 5090 32 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - gemma-2-9b-it on RTX-PRO-6000-Blackwell-Server: Fits, cost withheld (https://runinfra.ai/calc/gemma-2-9b/rtx-pro-6000-blackwell-server/50k-req-day): gemma-2-9b-it fits in memory on 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB, but the serving configuration is not established. gemma-2-9b-it fits 1 x NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB at 50,000 req/day. No complete throughput evidence can price this configuration. - gemma-2-9b-it on MI300X: At least $3,783.99/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/mi300x/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $3,783.99 using 2 minimum replicas of 1 x AMD Instinct MI300X at 50,000 requests/day. gemma-2-9b-it on 2 minimum replicas of 1 x AMD Instinct MI300X at 50,000 req/day: GPU rent at least $3,783.99/mo, a lower bound, not an estimate. - gemma-2-9b-it on MI325X: At least $2,775.90/mo at 50K req/day (https://runinfra.ai/calc/gemma-2-9b/mi325x/50k-req-day): The monthly GPU rent for gemma-2-9b-it is at least $2,775.90 using 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 requests/day. gemma-2-9b-it on 1 minimum replica of 1 x AMD Instinct MI325X at 50,000 req/day: GPU rent at least $2,775.90/mo, a lower bound, not an estimate. ## Use-case recipes - Voice agent technical reference (https://runinfra.ai/use-cases/voice-agent): reference material for this workload pattern. - AI assistant technical reference (https://runinfra.ai/use-cases/ai-assistant): reference material for this workload pattern. - Embeddings technical reference (https://runinfra.ai/use-cases/embeddings): reference material for this workload pattern. - RAG search technical reference (https://runinfra.ai/use-cases/rag-search): reference material for this workload pattern. - Document AI technical reference (https://runinfra.ai/use-cases/document-ai): reference material for this workload pattern. - Transcription technical reference (https://runinfra.ai/use-cases/transcription): reference material for this workload pattern. ## GPU reference - NVIDIA L4, 24 GB VRAM - NVIDIA L40S, 48 GB VRAM - NVIDIA A100, 80 GB VRAM - NVIDIA H100, 80 GB VRAM - NVIDIA H200, 141 GB VRAM - NVIDIA B200, 192 GB VRAM ## Newsroom articles --- URL: https://runinfra.ai/news/b200-beats-the-lpu Category: Engineering Published: August 6, 2026 Updated: August 9, 2026 Author: Jaber Jaber (Founder and researcher, RunInfra) Read time: 23 min read Tags: B200, gpt-oss-120b, SGLang, Speculative decoding, Cerebras, Groq, Benchmark # With software alone, one B200 beats the LPU and gets close to Cerebras Stock SGLang served gpt-oss-120b at 411 tokens per second on one rented B200. Four settings took the same card and the same weights to 1,366, and single prompts to 2,215. No kernel written, nothing recompiled. Correction, August 9, 2026: Artificial Analysis changed its provider medians after this article was published on August 6. On August 6 the comparison points used here read Cerebras 1,991 tok/s, SambaNova 708, Groq 476, Google Vertex 423, Databricks 324, Azure 300, Amazon 101, and CoreWeave 33. On August 9 the source reads Cerebras 1,891, SambaNova 706, Groq 477, Google Vertex 391, Azure 312, Baseten 195, Amazon 86, and CoreWeave 36; Databricks has no reported speed in the current chart. The current prose and tables below use the August 9 values. The original August 6 chart images remain visible as dated publication snapshots. Every year now somebody ships new silicon built for inference. Cerebras put the weights on a wafer, Groq built the LPU, SambaNova built the RDU, and all three of them are real machines that move tokens fast. I like that this is happening. Inference is worth designing hardware for. I spent the last two years studying NVIDIA and AMD GPU architecture to build a code editor for GPU kernels. Reading those manuals all day teaches you what a card can do, and then you look at what people are getting out of the same card and the two numbers do not match. So I stopped trying to write every kernel by hand and started building agents that write them. I open sourced AutoKernel, which generates and tunes GPU kernels, small and readable so you can actually follow what it does. Then AutoMegaKernel, which fuses a whole decode step into persistent kernels so the intermediates never leave the chip. All of it is for the same thing, getting models to run properly on the cards people already own. That is why the leaderboards bother me. When the board read 423 GPU tokens per second on August 6 for a model whose byte budget I could work out on paper, my first thought was not that the GPU was slow. My first thought was that nobody wrote the software. So I decided to check whether I could close that gap on a single B200 without touching anything but the software. As of August 9, Artificial Analysis lists Cerebras at 1,891 tokens per second. The fastest listed GPU provider is Google Vertex at 391, then Azure at 312 and Baseten at 195. That puts Cerebras 4.8x ahead of the fastest listed GPU, close to the 5x lead in Cerebras's published comparison (cerebras.ai/blog/blackwell-vs-cerebras). (Image: Cerebras blog text headed "Cerebras, still the fastest inference in 2025", saying Blackwell improves GPU inference by 2 to 3x and that Cerebras is the only architecture that outperforms NVIDIA, with a 5x lead on OpenAI's flagship open-weight model.. Cerebras, in their own words (cerebras.ai/blog/blackwell-vs-cerebras).) The 1,891 is a hardware answer and a good one. They keep the weights in SRAM instead of HBM, so the bottleneck the rest of us fight does not exist on their machine. I am not going to argue with that and I did not beat it. The board read 423 when I ran the comparison and reads 391 as of August 9. That GPU reference point is the number I went after. I rented a single B200 on Modal, which is what I reach for when I want a GPU without thinking about it, and started measuring. Out of the box I got 411 tokens per second. I changed four settings and the same card running the same weights gave me 1,366, and single prompts hit 2,215. I did not write a kernel and I did not recompile anything. - Stock SGLang: 411 tok/s. 26 percent of this card's memory bandwidth - After four settings: 1,366 tok/s. 72 percent, and 3.5x the fastest GPU provider - Best single prompt: 2,215 tok/s. above Cerebras, on one rentable card That is 3.5x past the fastest GPU provider on the current board. It clears Groq at 477 and SambaNova at 706, and it puts one rented card at 72 percent of the wafer. ## The abstraction lies to you, and it is not lying about the hardware Every layer between you and the chip is there to stop you thinking about the chip, and it works. You call generate, tokens come out at some rate, and that rate feels like the speed of the machine. It is not. It is the speed of whatever defaults you happened to call. That is why people blame the hardware when something is slow, and why they are almost always wrong. The framework cannot tell you that it is the thing holding you back. It throws no error, everything looks fine, and the number it gives you is steady and repeatable, so it feels like physics when it is a config file. The only way out is to work out what the card can do yourself, from bytes and bandwidth, and compare. That number does not care what framework you use. For gpt-oss-120b on a B200 the arithmetic is short. Decoding one token reads every active weight exactly once, which is 1.91 GB of attention in bf16, 1.90 GB of top-4 routed experts in mxfp4, and 1.16 GB of lm_head in bf16, so 4.97 GB per token. A B200 has 8 TB/s of HBM3e (nvidia.com/en-us/data-center/dgx-b200), so the floor is 0.621 milliseconds per token and the ceiling is about 1,610 tokens per second. Stock SGLang took 2.43 milliseconds per token. Seventy four percent of every token was not reading weights. (Image: Horizontal bar of one decoded token at stock settings, 2.43 milliseconds total. A red segment of 0.621 milliseconds is the weight read; the remaining 1.809 milliseconds is grey overhead.. Red is the irreducible read of 4.97 GB of weights at 8 TB/s, which caps this model at 1,610 tok/s. Grey is everything else.) The chip sat idle three quarters of the time. Every instinct that says buy the faster card is aimed at the 26 percent that was already working. ## What I measured, and how I ran everything on a single B200 on Modal with 178.4 GiB of HBM3e and 148 SMs. Model is openai/gpt-oss-120b, which ships natively in MXFP4 and activates about 5.1B of its 120B parameters per token. Serving through SGLang 0.5.16, greedy at temperature zero, single stream, one prompt at a time, four fixed prompts, three repetitions, median. The number reported is decode rate excluding prefill, which is how Artificial Analysis defines Output Speed, the average number of tokens received per second after the first token is received (artificialanalysis.ai/methodology). I picked this model because all three ASIC vendors publish on it, so the comparison is the same weights and the same metric. Provider | tok/s | Silicon --- | --- | --- Cerebras | 1,891 | wafer SambaNova | 706 | RDU Groq | 477 | LPU Google Vertex | 391 | GPU Azure | 312 | GPU Baseten | 195 | GPU Amazon | 86 | GPU CoreWeave | 36 | GPU (Image: Archived Artificial Analysis provider chart captured August 6, 2026 for gpt-oss-120b. Its publication-time values are retained as historical evidence and no longer match the live source.. Archived August 6 publication snapshot. The live August 9 source values are in the corrected table above; this image is retained so the original published comparison remains auditable.) Look at the measured GPU listings on the current board. They land 10.8x apart, from Google Vertex at 391 down to CoreWeave at 36, on the same open weights and metric. The public chart does not isolate GPU generation from settings, kernels, or serving code, so this is evidence that the delivered stack matters, not a clean 10.8x software-only estimate. ## The ladder Step | tok/s --- | --- stock defaults | 411.2 latency knobs | 429.9 n-gram speculative decoding | 628.5 speculative_attention_mode=decode | 697.1 wide n-gram search | 1,165.1 and 1,366.2 (Image: Bar chart captured for the August 6 publication: stock 411, plus knobs 430, plus n-gram 629, plus decode 697, and plus wide search 1,366 in red, against the publication-time Cerebras line at 1,991 and Groq line at 476.. The grey bars are the settings anyone can copy. The red one is where the n-gram search width lands it. Reference lines remain at the August 6 snapshot values; the August 9 medians are Cerebras 1,891 and Groq 477.) The first rung is small. SGLang defaults stream_interval to 1, so it detokenizes and dispatches on every single token, and scheduler_recv_interval to 1, so it polls every iteration. On a 0.621 millisecond budget that is real money, and fixing it bought 4.5 percent. The second rung is speculative decoding, the same trick that gave Groq its six times. A draft proposes several tokens, the target verifies them in one forward pass, and every accepted token is a token produced without a separate 4.97 GB weight read. I could not use a trained draft head, for reasons in the failures section, so this is n-gram speculation, which drafts by matching patterns in the text generated so far. The third rung was a flag I had never touched. speculative_attention_mode defaults to prefill, and setting it to decode gave 9.7 percent. The last row is two numbers because I ran that configuration twice and got both, a 17 percent spread on identical settings. N-gram acceptance depends on what the trie has built up, so the same config lands in a different place each run. If I quoted the better one I would be picking a number, not measuring one. ## The rung that mattered was found by breaking it I was told, reasonably, that the CPU n-gram proposer was serializing with the GPU and that its search cost was the bottleneck. So I cut it down, max_bfs_breadth from 10 to 2 and max_trie_depth from 18 to 8, expecting overhead to fall. Throughput collapsed to 433.0. The failure was worth more than a small win: It says the search is not overhead, it is the thing that earns the tokens. Every extra candidate the proposer explores is a chance to accept another token, and every accepted token skips a full 4.97 GB read. Spending a few hundred microseconds of CPU to avoid 0.621 milliseconds of HBM traffic is a trade you want to make constantly. So I ran the knob the other way, and the curve has a sharp peak. max_bfs_breadth | tok/s --- | --- 2 | 433.0 10, the default | 697.1 24 | 1,165.1 32 | 1,110.4 48 | 341.1 (Image: Line chart of tokens per second against n-gram search breadth: 433 at breadth 2, 697 at the default 10, a peak of 1,165 at 24, 1,110 at 32, and a collapse to 341 at 48.. Wider search costs CPU and buys accepted tokens, and each accepted token skips a 4.97 GB read. Past 24 the trade stops paying.) The shipped default is 10, and you can read it straight off the documentation. (Image: SGLang speculative decoding documentation table listing the n-gram parameters, with maximum BFS breadth defaulting to 10.. The n-gram parameters from the SGLang speculative decoding docs (docs.sglang.io/docs/advanced_features/speculative_decoding). Maximum BFS breadth, default 10.) For single-stream decode that leaves about 40 percent on the floor, and past 24 the CPU search stops paying for itself and falls off a cliff. The default is not a mistake. It is set for throughput serving, where many requests share the GPU and CPU time is tight. It is the wrong default for one user waiting on one stream, and nothing in the stack tells you which of the two you are doing. ## Where that lands Provider | tok/s | Ratio --- | --- | --- one B200, two runs | 1,165.1 to 1,366.2 | Groq | 477 | 2.44x to 2.87x faster SambaNova | 706 | 1.65x to 1.94x faster Cerebras | 1,891 | 0.62x to 0.72x Per prompt, the two runs measured 447.8, 1930.8, 1882.5, 219.1 and then 525.3, 2207.0, 2215.6, 291.9 tokens per second. Three of those eight readings are above Cerebras's current 1,891 median, on one rentable GPU. (Image: Bar chart of four prompts under one configuration: open prose lowest, technical explanation next, then code and structured reasoning both above 2,200 tokens per second, with reference lines for Groq and Cerebras.. Acceptance tracks predictability. The reference lines remain at the August 6 snapshot values; against the August 9 Cerebras median of 1,891, three of the eight measured prompts are higher.) That spread is the most useful thing in the post, so I am showing it instead of hiding behind a median. N-gram speculation drafts by matching text it has already produced, so structured reasoning where phrasing repeats runs above 2,200 while open-ended prose that never repeats itself runs at 219. The trick pays off exactly as much as the output is predictable. A vendor reporting one number is averaging over their own prompt mix, and you cannot see this shape at all. ## What speculation does and does not change Greedy verification is what keeps this honest. A drafted token is accepted only where it matches the token the target model would have produced on its own, and on the first mismatch the rest of the draft is thrown away and the target's token is taken instead. So speculation under greedy verification is exact, not approximate. It changes how many forward passes you spend, not which tokens come out. If you go and check this yourself: Comparing a speculative run against a differently configured non-speculative run does not test speculation, it tests the kernels. Switching the attention backend changes the order of floating point reductions, and wherever two logits are nearly tied that reorder flips the argmax. Your control has to hold every other setting fixed and flip speculation on its own. And compare token ids, never text, with an explicit length check rather than zip, which truncates to the shorter side and will report a run that died halfway as bit exact. I know because I wrote that bug earlier in this project and it passed a kernel that had crashed. ## What did not work, which is most of it Fifteen configurations, and the failures tell you more than the wins. Both published EAGLE-3 draft heads scored below running with no speculation at all, 417.6 for NVIDIA's and 325.0 for the SGLang team's, because acceptance was near zero. NVIDIA trained theirs against their NVFP4 checkpoint and this is the MXFP4 one, and a draft that never guesses right costs you the draft pass for nothing. Getting NVIDIA's head to load at all is worth recording. It failed with a tensor mismatch, 8640 against 5760, which is three times the hidden size against two. SGLang v0.5.16 maps the draft's requested capture layers with an off-by-one, so the last of three requested layers falls outside the loop, and the same version silently drops the checkpoint's input_norm weight because it never builds that module. Both were fixed upstream in commit 5df193b4ac, merged six days after v0.5.16 was tagged. A nightly with the fix loaded it, and it still lost. Every attention backend except triton is unavailable here. FlashAttention 3 requires SM 80 to 90 and a B200 is SM100. FA4 forces a page size of 128 and trtllm_mha forces 64, and SGLang then refuses to run a speculative tree wider than one token on a paged backend because it produces incorrect results. I was glad to hit that guard instead of quietly getting wrong answers. The triton MoE runner OOMs at 156 GiB because it dequantizes MXFP4 back to bf16. torch.compile asserts and then OOMs on the same path. DFLASH with a real draft model reached 493.5, better than nothing and worse than n-gram. ## The honest caveats They are serving real traffic on production endpoints and I ran four fixed prompts on one rented box, so this is not a clean head to head and I am not going to pretend it is. Artificial Analysis also feeds 10,000 input tokens and my prompts are about thirty, and a longer prompt means more KV cache to read on every step, so they are doing the harder job. That helps my number, so knock it down a bit when you compare. A wafer holding weights in SRAM is a genuine architectural advantage on this problem and not marketing. Cerebras is fast because on-chip memory removes the bottleneck I have been describing, and no amount of kernel work turns HBM into SRAM. Their median still beats mine. A card you can rent by the hour running weights you can download reaches 1,165 to 1,366 tokens per second after four settings, and that beats two of the three custom chips. The scripts are in the repo and so is the raw per prompt JSON, so you can go and look at the spread yourself instead of trusting me. ## Why I build what I build Stock configuration reached 26 percent of this machine's memory bandwidth. Four settings took it to 72 percent. Nobody replaced the hardware, and the 2.8x was sitting inside an abstraction that reported no error the entire time. I am young and I am new to kernels and I am not the best person in the world at writing them. What I am good at is building agents that write them, and I think that is the more useful skill. The people who can hand tune a kernel are rare and they do not scale, and every new model on every new card needs the work done again. Kernel work is where the next decade of performance comes from, not the next fab. The code editor came out of those two years reading architecture manuals, and I have stopped supporting it. AutoKernel and AutoMegaKernel came next, and both of them exist to do by machine what I did by hand in this post. All of it is about making models run on the cards people own, not the ones in press releases. I built RunInfra so you do not have to do any of this by hand. You point it at any model on Hugging Face and at whatever cards you already have. It works out your ceiling, writes and tunes the kernels that close the gap, and checks the output is bit exact before it ships anything. The work I did by hand here it does on its own. (Image: The RunInfra composer, headed Optimize any Hugging Face model for production, with the models going in listed on the left and the GPU each one was measured on coming out on the right.. runinfra.ai) The order of operations matters and the industry keeps getting it backwards. Before designing a different architecture we should squeeze what is already racked, because most of that silicon runs at a fraction of what it can do and the missing part is code nobody has written. Taping out a chip to solve a problem you have not first solved in software means paying eighteen months and a fab run for something a configuration flag might have handed you the same week, and the flags I changed here were worth 3.3x. I am not taking anything away from Cerebras. Keeping the weights in SRAM kills the exact bottleneck I have been describing, that is a real fix for a real problem, and they are still ahead of me. But a lot of what looks like a hardware gap on a leaderboard is software somebody has not written yet, and you only find that out by working out the ceiling yourself and noticing you are at 26 percent of it. The thing people call hardware design is mostly translation anyway. A model is a graph of operations and a chip is a set of execution units and memory levels. Getting from one to the other is a compiler and kernel problem end to end. You are deciding what fuses, what stays in registers, what spills to HBM and which loop order the tensor cores want. That translation layer is where the performance lives, it is software, and it is the same layer whether the silicon underneath is from NVIDIA or AMD or something taped out last year. Which is why I do not think any of this is permanent, and the clearest case is AMD. The silicon is good and the bandwidth is there. The gap to NVIDIA on real work is the stack, not the transistors. If software is the larger share then the gap is closable by writing software, which is a far better position than needing a new fab. We are scaling our agents to AMD next, because the interesting test is not making a fast chip faster, it is reaching NVIDIA class numbers on hardware people have written off. It is not just the chip companies either. OpenAI and Broadcom put out an inference chip called Jalapeno in June, and Anthropic confirmed an in house silicon team in August, both of them talking about cutting per token cost roughly in half (openai.com/index/openai-broadcom-jalapeno-inference-chip, techtimes.com/articles/323238). Two of the best software companies in the world decided the next place to spend is the hardware layer. I read that differently to most people. If you are getting half your inference cost back by taping out a chip, part of what you are really buying back is the efficiency you never got out of the GPUs you already have. I do not think that is a criticism of them, it is just the same 26 percent I measured in this post, sitting at a scale where it is worth a fab run to fix. What I am less sure about is what happens next. CUDA is the reason a B200 is easy to reach 72 percent on, and CUDA is eighteen years of compilers, libraries, kernels, profilers and people who know where the bodies are buried. That is the moat, not the transistors. A new chip starts at zero on all of it. It can be a better design and still lose, because what decides how much of a chip you can actually use is the software around it, and that takes years to build. So the question I would ask about any custom chip is not how fast the silicon is. It is whether the software keeps up. Will it still be fast on the model you switch to next quarter, on a trick nobody has invented yet, on a shape the compiler was never written for. A GPU answers all of that with a recompile. Custom inference chips are a bet that one workload stays still long enough to bake it into silicon. Speculative decoding did not exist in its current form three years ago and the next trick does not exist yet either. When it arrives, the people on programmable hardware with a mature compiler will ship it in weeks. The people who taped it out will not. I remember when I moved from calling PyTorch functions to writing the kernels underneath them, I thought I was going one level down for performance, and I am more sure now that I was going one level down for control, because performance is what you get when you stop accepting whatever the default decided on your behalf. So when people ask me why I keep working on this instead of something with a nicer demo, that is the answer. The hardware is going to keep being good. The part that decides whether you see any of it is the software, and right now almost nobody is writing it. That is the piece that matters, and it is the piece I want anyone to be able to reach for. If you have GPUs sitting at a fraction of what they can do: That is what RunInfra is for. Any model on Hugging Face, any hardware you already own, kernels written and tuned for your setup, and every one of them checked for bit exact output before it ships. https://runinfra.ai ## Answer summary A single rented NVIDIA B200 served openai/gpt-oss-120b at 411 tokens per second on stock SGLang 0.5.16 and at 1,165 to 1,366 tokens per second after four configuration changes, with individual prompts reaching 2,215. The changes were stream_interval and scheduler_recv_interval, n-gram speculative decoding, speculative_attention_mode=decode, and widening speculative_ngram_max_bfs_breadth from the default 10 to 24. No kernel was written and nothing was recompiled. The arithmetic ceiling for this model on this card is about 1,610 tokens per second, from 4.97 GB of weights read per token against 8 TB/s of HBM3e, so stock configuration was using 26 percent of the machine's memory bandwidth and the tuned configuration 72 percent. ## Frequently asked Q: How fast can one B200 serve gpt-oss-120b with software changes alone? A: 411 tokens per second on stock SGLang 0.5.16, and 1,165 to 1,366 tokens per second after four configuration changes on the same card with the same weights. Individual prompts reached 2,215. No kernel was written and nothing was recompiled. Q: Does a GPU beat Groq or Cerebras on gpt-oss-120b? A: It beats Groq and SambaNova and does not beat Cerebras. One tuned B200 measured 1,165 to 1,366 tokens per second against Groq at 477 and SambaNova at 706, which is 2.44x to 2.87x faster than Groq. Cerebras publishes 1,891, so the tuned GPU reaches about 62 to 72 percent of it. Q: What is the ceiling for gpt-oss-120b on a B200? A: About 1,610 tokens per second for single-stream decode. Decoding one token reads every active weight once, which is 1.91 GB of attention in bf16, 1.90 GB of top-4 routed experts in mxfp4 and 1.16 GB of lm_head in bf16, so 4.97 GB per token. Against 8 TB/s of HBM3e that is a floor of 0.621 milliseconds per token. Q: Which SGLang settings mattered most for single-stream decode? A: Widening the n-gram search was worth the most by far. speculative_ngram_max_bfs_breadth ships at 10, and raising it to 24 took the same configuration from 697 to 1,165 tokens per second. Past 24 the CPU search stops paying for itself and throughput falls off a cliff, to 341 at 48. Q: Does speculative decoding change the output? A: Not under greedy verification. A drafted token is accepted only where it matches the token the target model would have produced, and on the first mismatch the remaining draft is discarded and the target's token is used. Speculation changes how many forward passes you spend, not which tokens come out. Q: Why do the same settings give different numbers per run? A: N-gram speculation drafts by matching text it has already produced, so acceptance depends on how predictable the output is and on what the trie has built up. The same configuration measured 1,165.1 and 1,366.2 across two runs, and per prompt ranged from 219 on open-ended prose to 2,215 on structured reasoning. --- URL: https://runinfra.ai/news/lossless-inference Category: Research note Published: August 3, 2026 Author: Jaber Jaber (Founder and researcher, RunInfra) Read time: 12 min read Tags: LLM inference, Quantization, Speculative decoding, FlashAttention, KV cache, Serving # Lossless Inference How to make LLM serving faster without touching the model. Exact kernels, speculative decoding, lossless compression, KV reuse and scheduling, with the math and how to verify it with logit parity. Quantization became the default way to speed up inference because it is the easiest one. In vLLM it is one flag or one swapped checkpoint (docs.vllm.ai), while doing the real optimization work like writing better kernels takes months. So people swap the checkpoint and move on, and the same people who would never ship an unreviewed prompt change end up shipping a quantized model without running a single eval, because the flag makes it feel like an infra setting when it is actually a model change (Image: fig1. Figure 1: panel (a) is the gap, benchmark score stays roughly flat across precision while long horizon task success falls. Panel (b) is where the loss lands, one flipped decision mid run and everything after inherits the error. Illustrative) Quantization is a trade, not an optimization. You save memory and you pay for it in quality, and the cost is hard to see because post training quantization moves every weight a little while the standard benchmarks barely react, since they are saturated and they test one step at a time. The real loss shows up in long reasoning chains and agentic coding runs where hundreds of dependent decisions stack on each other, so the benchmark tells you the two models are the same while your agent's success rate tells you they are not. This post is about the serving stack that gets the speed without paying that cost ## The compounding math If quantization noise flips one decision per step with probability p, then over n dependent steps the task survives with probability (1-p)^n. A 2 percent hit per step sounds free until you run 20 steps, because 0.98^20 = 0.67 and a third of your tasks now fail, and at 50 steps you are down to 0.36 If you hold p at half a percent instead, survival over 20 steps stays at 0.90, and that is the difference between an agent you trust and an agent you babysit. The model never got visibly dumber, it just got 2 percent noisier per decision and the compounding did the rest (Image: fig2. Figure 2: task survival (1-p)^n against step count for four values of p. The marked points are 0.98^20 = 0.67 and 0.98^50 = 0.36. The curves are exact) ## What lossless means Lossless inference means the stack gets faster while the model stays exactly the same model, and I split it into two tiers Bit exact means the optimized stack returns the same logits as the reference on every input, f_opt(x) = f_ref(x). Kernels and fusion, compilation, KV reuse, scheduling and lossless weight compression all live in this tier because they change how the math runs without changing the math itself Distribution exact means the optimized stack samples from the same distribution as the reference, P_opt(y | x) = P_ref(y | x). Speculative decoding with rejection sampling lives here, where single samples can differ but the distribution never does Everything else, PTQ, pruning and distillation, changes the function itself, which means you are serving a different model under the same name. This definition matters because it moves the burden of proof, since a lossless change needs no eval run when the outputs are the evidence, while a lossy change needs a full eval on your workload and not just on MMLU. Bit exact also composes, because a pipeline of bit exact stages is still bit exact, so I can stack five of them without running five eval cycles (Image: fig3. Figure 3: the definition as a diagram. Two lossless tiers with their guarantees and the techniques inside each, while lossy methods edit the function and pay the compounding cost) ## The lossless stack, bottom up The stack starts with kernels. Inference at serving batch sizes is memory bandwidth bound, since an H100 SXM moves about 3.35 TB/s from HBM while doing about 989 dense BF16 TFLOP/s (nvidia.com/en-us/data-center/h100/), so a kernel earns its speed by moving fewer bytes rather than by doing less math. FlashAttention proved the point years ago by computing exact attention with reordered memory traffic and getting big speedups with identical outputs (arxiv.org/abs/2205.14135), and the reordering is valid because online softmax computes the same normalization from running statistics, so nothing about the result changes (Image: flashattention. The FlashAttention idea in one picture. Tile the computation to fit SRAM, never materialize the full attention matrix, and get the exact same output faster. Diagram from the official FlashAttention repository (Dao et al.), BSD 3 Clause license: github.com/Dao-AILab/flash-attention) Fusion pushes the same idea further by keeping intermediate values in registers and shared memory instead of round tripping through HBM, and a megakernel takes it to the limit by running the whole decode step as one launch. I built this myself. On the same H100 my engine bonsai-turbo runs 1.76x the vendor's own llama.cpp fork, going from 85.5 to 151.1 tokens per second with the same outputs (github.com/RightNow-AI/bonsai-turbo), and my kernel search system AutoKernel reached 1.31x over cuBLAS on its best kernels across 95 experiments spanning 18 to 187 TFLOPS while holding first place on the B200 vectorsum leaderboard (github.com/RightNow-AI/autokernel) Speculative decoding comes next. A small draft model proposes g tokens, the target model verifies them in one pass, and a modified rejection sampling step keeps the output distribution exactly the target's, which was proved independently twice in 2023 (arxiv.org/abs/2211.17192, arxiv.org/abs/2302.01318). The expected number of tokens per target pass is (1-a^(g+1))/(1-a) at acceptance rate a, so draft accuracy is everything, and the EAGLE line is what drove it up (arxiv.org/abs/2401.15077). EAGLE-2 states it plainly ("the distribution of the generated text remains unchanged") and reports 3.05x to 4.26x (arxiv.org/abs/2406.16858), while EAGLE-3 reports up to 6.5x (arxiv.org/abs/2503.01840) (Image: eagle draft tree. The EAGLE method. A small head extrapolates features from the target model, drafts a token tree, and the target verifies the whole tree in one pass. Diagram from the official EAGLE repository (SafeAILab), Apache 2.0 license: github.com/SafeAILab/EAGLE) Then there is lossless weight compression. BF16 weights carry less than 16 bits of real information per parameter, so DFloat11 entropy codes them down to about 11 bits, which lands at about 70 percent of the original size with outputs that stay bit for bit identical to the original model (arxiv.org/abs/2504.11651). That covers most of the memory saving that quantization promises without the quality bill KV reuse is the next layer. Paged attention ended KV cache fragmentation and raised real batch sizes (arxiv.org/abs/2309.06180), and prefix caching with radix trees makes shared history get computed once (arxiv.org/abs/2312.07104). This is safe because cache entries are a pure function of the prefix, so reusing them at the same positions cannot move the logits (Image: radixattention. RadixAttention step by step. The KV cache becomes a radix tree where chat turns and branches share every common prefix and cold nodes get evicted. Diagram from the SGLang blog repository (LMSYS), MIT license: github.com/lm-sys/lm-sys.github.io) Scheduling sits on top. Continuous batching refills the batch at token boundaries instead of waiting for the longest request to finish (usenix.org/conference/osdi22/presentation/yu), chunked prefill stops a long prompt from stalling everyone else's decode (arxiv.org/abs/2308.16369), and prefill decode disaggregation puts the compute bound half and the bandwidth bound half of the workload on different machines (arxiv.org/abs/2401.09670). None of these touch a single logit (Image: fig4. Figure 4: panel (a) is the serving pipeline with every lossless technique in place, where BE marks bit exact stages and DE marks distribution exact ones. Panel (b) is the H100 SXM roofline from public specs with the ridge at 295 FLOP per byte. Panel (c) is the expected tokens per target pass under speculative decoding. Panels (b) and (c) follow directly from the stated specs and formula) ## The moving line The lossless boundary is drawn at the shipped checkpoint, and the checkpoint itself is moving. Kimi K2 Thinking shipped with native INT4 weights because Moonshot ran quantization aware training in post training, put INT4 on the MoE weights and reported every official benchmark at INT4 with roughly 2x faster generation (huggingface.co/moonshotai/Kimi-K2-Thinking). OpenAI shipped gpt-oss with MXFP4 MoE weights at 4.25 bits per parameter, which covers over 90 percent of the parameters and is the reason gpt-oss-120b fits on one 80 GB GPU (arxiv.org/abs/2508.10925) For these models serving INT4 or MXFP4 is not lossy quantization, it is the checkpoint. The reference you must not degrade is already low precision, and the vendor already paid the quality cost during training where it belongs, so quantization is turning into a training decision while squeezing a BF16 checkpoint after the fact stays a serving hack. My test is simple, if the weights you serve are the weights the vendor benchmarked then you are serving the model, and if you squeezed them afterward then you are serving your own edit of the model (Image: fig5. Figure 5: where precision is decided against what it costs relative to the shipped checkpoint. The top left cell is where the big model releases moved in 2025, and the timeline marks the shift) ## Verification Lossless is a claim you can test, and I test it with logit parity. You run the same prompts through the reference stack and through yours with greedy decoding, then compare the maximum absolute logit difference per token, where zero is the pass bar for a bit exact claim and a distribution test is the bar for speculative decoding. I published this for bonsai-turbo, where 32 of 32 prompts pass against the vendor fork across all four engine configurations (github.com/RightNow-AI/bonsai-turbo). When parity fails the diff shows me where to look, and it is usually a fused kernel that cut a corner, a RoPE mismatch or a sampler that renormalizes differently Moonshot now runs the same kind of check on everyone who serves its model. The K2 Vendor Verifier sends 4,000 identical requests to every third party K2 provider and scores them against the official API on tool call trigger F1 and schema accuracy (github.com/MoonshotAI/K2-Vendor-Verifier), and the spread was real, with schema accuracy running from 100 percent at the best providers down to about 73 to 76 percent on stock open source engines before fixes landed. Model quality became a serving stack property and model vendors have started measuring it Floating point addition is not associative, so the reduction order inside a kernel changes the low bits, and that order changes with batch size. Thinking Machines measured the consequence when 1,000 temperature zero completions on stock vLLM gave 80 distinct outputs, while batch invariant kernels brought all 1,000 back bitwise identical at 1.6x to 2x the runtime (thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/), and Chen et al. scoped their speculative decoding proof "within hardware numerics" for the same reason (arxiv.org/abs/2302.01318). So lossless means the stack adds no error beyond the numerics you already accepted, and it does not mean bitwise identical outputs across batch sizes unless you also pay for batch invariance (Image: fig6. Figure 6: panel (a) is the parity protocol I run, one prompt set through both stacks, diff the logits, zero or go debug. Panel (b) is what the diff looks like, where a lossless stack sits exactly at zero. The nonzero values are illustrative) ## Why lossless wins The economics settle the argument. Quantization saves a fixed factor once and keeps paying quality on every step of every task, while lossless wins multiply with each other, kernels times speculative decoding times KV reuse times scheduling, with every term keeping the model exact. One term alone already shows what lossless buys on real hardware: (Image: eagle3 speedup. Measured lossless speedups from the EAGLE line, up to 5.6x on Vicuna 13B and 5.0x on DeepSeek R1 LLaMA 8B with the output distribution unchanged. Chart from the official EAGLE repository (SafeAILab), Apache 2.0 license. Not my benchmark but a real one) (Image: fig7. Figure 7: lossless wins compound across the stack. The bar sizes are illustrative ranges, not a benchmark) I remember when I started working on this, I thought lossless would be the future, and I am more sure of it now. At RunInfra I build exactly this, serving open models on the full lossless stack and verifying it with logit parity against the reference ## Frequently asked Q: What does lossless inference mean? A: Bit exact means the optimized stack returns the same logits as the reference on every input, f_opt(x) = f_ref(x). Kernels and fusion, compilation, KV reuse, scheduling and lossless weight compression all live in this tier because they change how the math runs without changing the math itself Q: Why does quantization hurt long agentic runs? A: If you hold p at half a percent instead, survival over 20 steps stays at 0.90, and that is the difference between an agent you trust and an agent you babysit. The model never got visibly dumber, it just got 2 percent noisier per decision and the compounding did the rest Q: How do you verify a lossless claim? A: Lossless is a claim you can test, and I test it with logit parity. You run the same prompts through the reference stack and through yours with greedy decoding, then compare the maximum absolute logit difference per token, where zero is the pass bar for a bit exact claim and a distribution test is the bar for speculative decoding. I published this for bonsai-turbo, where 32 of 32 prompts pass against the vendor fork across all four engine configurations (github.com/RightNow-AI/bonsai-turbo). When parity fails the diff shows me where to look, and it is usually a fused kernel that cut a corner, a RoPE mismatch or a sampler that renormalizes differently --- URL: https://runinfra.ai/news/vllm-vs-sglang-vs-tensorrt-llm-benchmark Category: Engineering Published: June 20, 2026 Author: Jaber Jaber (Founder and researcher, RunInfra) Read time: 11 min read Tags: vLLM, SGLang, TensorRT-LLM, Benchmark, LLM inference, GPU # vLLM vs SGLang vs TensorRT-LLM: a reproducible benchmark We benchmarked the three main LLM inference engines on the same model, GPU, and request stream, and open-sourced the harness. The winner flips with how you load it. Scoped to Llama-3.1-8B on H100 and L40S. ## Key takeaways - There is no single fastest engine. The winner changes with the operating point. - TensorRT-LLM had the lowest first-token latency at moderate load (235 ms vs vLLM 514 ms at concurrency 32). - vLLM reached the highest throughput at saturation (5,333 tok/s on one H100) and the lowest cost. - fp8 is the cheapest way to serve this model: vLLM at fp8 is $0.158 per 1M tokens. - The cheap L40S GPU costs about twice as much per token as the H100. Everything below is scoped to what we tested: Llama-3.1-8B-Instruct, one H100 80GB and one L40S 48GB, bf16 and fp8. We do not generalize past it. Versions, as of June 20 2026: vLLM 0.23.0, SGLang 0.5.13, TensorRT-LLM 1.2.1. The harness and every raw number are open source. ## The short version There is no single fastest engine. The winner changes with how you load it. If you care about | Use | On this benchmark --- | --- | --- Lowest first-token latency at moderate load | TensorRT-LLM | 235 ms TTFT at c32 vs vLLM 514 ms Highest throughput at saturation | vLLM | 5,333 tok/s on one H100 Lowest cost per token | vLLM at fp8 | $0.158 per 1M tokens Shared-prefix caching | a big win for both | throughput ~2.7x at 90% hit rate The rule for these configs: TensorRT-LLM for latency-bound serving below its throughput ceiling, vLLM for the highest throughput and the lowest cost. SGLang tracked vLLM within a few percent. Two things we are explicit about: We benchmarked TensorRT-LLM on its PyTorch backend, which has no ahead-of-time engine compile, so its peak throughput here is not the ceiling a compiled TensorRT engine would reach. And our single-prefix cache test is not the many-distinct-prefix case SGLang's RadixAttention is designed for, so we ran that separately too. We did not let either feed the headline. We ran every engine on the same weights, same GPU, and same request stream, and we open-sourced the harness so you can check us. The one clear surprise that holds: the cheap L40S GPU costs about twice as much per token as the H100. ## Why we ran this Hugging Face put TGI into maintenance mode in March 2026, so a lot of teams are re-choosing a serving engine right now. The comparisons you can find are point-in-time snapshots on someone else's workload. So we ran it ourselves, on our own fleet, and published the harness. If our numbers do not match yours, open the repo and find out why. ## How we measured One client drives all three engines with identical request streams, so the timing definitions are the same everywhere. TTFT is the time to the first streamed token. Throughput is output tokens per second. We warm up, then time three repeats per operating point and chart the mean; the 95 percent confidence intervals are tight and live in the committed raw data. Prefix caching is off for the unique-prompt sweeps so nothing gets a free cache hit. ## Throughput vs concurrency We swept concurrency from 1 to 256 at 1,024 input and 256 output tokens. At batch 1 the three engines are within four percent of each other. As load rises, TensorRT-LLM is competitive or best through concurrency 64, then flattens. vLLM keeps scaling and reaches the highest peak, 5,333 tok/s. (Chart: Output throughput vs concurrency, Llama-3.1-8B bf16, one H100) c | vLLM | SGLang | TensorRT-LLM --- | --- | --- | --- 1 | 158 | 153 | 152 8 | 1096 | 1005 | 1027 32 | 2921 | 2730 | 2978 64 | 4040 | 3898 | 4029 128 | 4943 | 4816 | 4774 256 | 5333 | 5235 | 4813 ## The latency vs throughput frontier The throughput table hides the more useful picture. Plot first-token latency against throughput and the tradeoff is clear. TensorRT-LLM holds the lowest TTFT across the whole mid range and keeps that edge up to about 4,000 tok/s, then hits its ceiling and TTFT climbs steeply. vLLM owns the high-throughput end: it reaches 5,333 tok/s at a lower TTFT than the other two get near that rate. (Chart: Latency vs throughput frontier (each point is a concurrency level)) vLLM: (158, 35), (1096, 177), (2921, 514), (4040, 790), (4943, 1658), (5333, 2221) SGLang: (153, 41), (1005, 221), (2730, 597), (3898, 997), (4816, 1761), (5235, 2543) TensorRT-LLM: (152, 32), (1027, 56), (2978, 235), (4029, 504), (4774, 1691), (4813, 2848) So the frontier has two owners. TensorRT-LLM for latency-sensitive serving below its ceiling, vLLM for maximum throughput. That is the whole point: the right engine is a function of your operating point, not a constant. ## TensorRT-LLM: we benchmarked the no-compile path Read this before you weigh the throughput numbers. We ran TensorRT-LLM on its PyTorch backend, which loads a Hugging Face checkpoint and serves it with no ahead-of-time engine build. That is the simplest path and the NVIDIA-recommended default, and it is what most people reach for first. It is not the maximum-throughput path. A compiled TensorRT engine, built offline with trtllm-build, chases peak tokens per second, at the cost of a multi-minute, GPU-architecture-locked compile step. We did not build it. So the 4,813 tok/s we measured is the PyTorch-backend peak, not TensorRT-LLM's ceiling, and we got no compile-time number. We make no claim that TensorRT-LLM cannot win raw throughput. What we can say: its PyTorch backend had the lowest first-token latency of the three across the mid range, and among the configs we ran, vLLM reached the highest throughput. ## Cost per token We turn throughput into dollars with one formula: cost per 1M tokens equals the GPU hourly price divided by sustained tokens per hour. The H100 is $3.95 an hour and the L40S is $1.95 an hour on Modal, observed June 20 2026. (Chart: Cost per 1M output tokens at each config's saturation (lower is better)) engine | bf16 / H100 | fp8 / H100 | bf16 / L40S --- | --- | --- | --- vLLM | 0.206 | 0.158 | 0.425 SGLang | 0.21 | 0.17 | 0.438 TensorRT-LLM | 0.228 | | Two things stand out. fp8 is the cheapest way to serve this model: vLLM at fp8 costs $0.158 per 1M tokens, 23 percent less than bf16, and it is faster too. The second is the one to remember. The L40S is half the hourly price of the H100, so it looks like the budget option. It is not. The L40S is about four times slower on this model, so it costs roughly twice as much per token, $0.425 versus $0.206. For Llama-3.1-8B at load, the H100 is both faster and cheaper per token. The cheap GPU is the expensive choice. ## Shared-prefix workloads RAG and agent traffic reuse a long system prompt across requests, so a prefix cache should help. We ran this two ways, because the shape of the reuse matters. ## One shared prefix First the simple case: one 2,048-token prefix sent to a fraction of requests, with that fraction varied from 0 to 90 percent, caching on and off. Both caches work. With caching off, throughput is flat. With caching on, it climbs about 2.7x as the hit rate reaches 90 percent. (Chart: Single shared prefix: throughput vs hit rate, caching on vs off (c32)) hit | vLLM (cache on) | vLLM (cache off) | SGLang (cache on) | SGLang (cache off) --- | --- | --- | --- | --- 0 | 999 | 910 | 937 | 879 50 | 1559 | 912 | 1528 | 873 90 | 2855 | 910 | 2488 | 875 vLLM and SGLang were within a few percent, vLLM slightly ahead. We do not read a winner into that gap. A single shared prefix is the easy case that any block-level cache handles well. It is not the workload SGLang's RadixAttention is built for. ## Many distinct prefixes, the RadixAttention case RadixAttention is designed for many distinct prefixes held in a tree at once, the shape you get from branching conversations and a pool of system prompts. So we ran that: 256 requests spread across a growing number of distinct 2,048-token prefixes, 8 then 32 then 128 of them, both engines with caching on. (Chart: Many distinct prefixes: throughput vs number of prefixes (c32, both caches on)) groups | vLLM | SGLang --- | --- | --- 8 | 2677 | 2493 32 | 2277 | 2086 128 | 1467 | 1376 RadixAttention worked. SGLang's logs show the prefixes served straight from its tree cache. But vLLM's prefix cache stayed ahead on throughput by 6 to 8 percent at every prefix count, with lower TTFT. So on Llama-3.1-8B on one H100 we did not reproduce a RadixAttention throughput win, even on the workload it targets. Both engines cache shared prefixes well; vLLM was a few percent faster, in line with its general edge on this hardware. RadixAttention may pull ahead with longer prefixes, deeper trees, or heavier eviction pressure than we tested. On our test it did not, and we are not going to claim otherwise. ## Reproduce it yourself The harness is open source under Apache-2.0 and every raw CSV is committed. One command runs the matrix under a hard cost cap; another re-plots from committed data with no GPU. Found a result that disagrees with yours? Open an issue with your run id and row. Corrections become a documented caveat or a v2. If you would rather not run this matrix for every model you ship, that is what RunInfra does. Point it at any Hugging Face model and it benchmarks the engines, picks the config for your latency and cost target, and deploys it serverless. ## FAQ vLLM or SGLang?: On plain unique-prompt traffic they are within a few percent, with vLLM slightly ahead on throughput, latency, and cost in our tests. Use vLLM as the default; reach for SGLang when its programmable frontend fits your workload. Is TensorRT-LLM worth it?: For latency-sensitive serving below its throughput ceiling, yes. It held the lowest time to first token across the mid range, 235 ms at concurrency 32 versus vLLM's 514 ms. We ran the PyTorch backend, which skips the engine compile. Which GPU should I use?: For Llama-3.1-8B at load, the H100. It is twice the hourly price of the L40S but about four times faster, so it costs roughly half as much per token. The L40S looks cheaper and is not. ## Method notes and caveats - One H100 80GB and one L40S 48GB on Modal, single GPU, tensor-parallel size 1. - Same Llama-3.1-8B-Instruct weights and same request stream across engines. Three timed repeats; the 95 percent confidence intervals are tight and live in the committed raw data. - Latency here is TTFT and per-request percentiles, not goodput at a fixed SLO. Whether 514 ms first-token at high load is acceptable depends on your use case; an SLO-goodput sweep is future work. - Numbers are self-reported on our harness. We do not cross-calibrate against an external suite like MLPerf, so the open harness is the check: re-run it and compare. - TensorRT-LLM ran the PyTorch backend (no engine compile), so there is no compile time to report and its peak number may differ from a built TensorRT engine. - The L40S sweep stopped at concurrency 128 (its saturation), the H100 at 256. - Engines move weekly. These numbers are vLLM 0.23.0, SGLang 0.5.13, TensorRT-LLM 1.2.1, as of June 20 2026. ## Answer summary A reproducible benchmark of vLLM 0.23.0, SGLang 0.5.13, and TensorRT-LLM 1.2.1 on Llama-3.1-8B-Instruct, one H100 and one L40S, bf16 and fp8. TensorRT-LLM had the lowest first-token latency at moderate load, vLLM the highest throughput at saturation and the lowest cost per token, and the L40S cost about twice as much per token as the H100. The harness is open source. ## Frequently asked Q: vLLM or SGLang? A: On plain unique-prompt traffic they are within a few percent, with vLLM slightly ahead on throughput, latency, and cost in our tests. Use vLLM as the default; reach for SGLang when its programmable frontend fits your workload. Q: Is TensorRT-LLM worth it? A: For latency-sensitive serving below its throughput ceiling, yes. It held the lowest time to first token across the mid range, 235 ms at concurrency 32 versus vLLM's 514 ms. We ran the PyTorch backend, which skips the engine compile. Q: What happened to TGI? A: Hugging Face moved Text Generation Inference to maintenance mode in March 2026. We note it for context and did not benchmark it as a contender. Q: Which GPU should I use? A: For Llama-3.1-8B at load, the H100. It is twice the hourly price of the L40S but about four times faster, so it costs roughly half as much per token. The L40S looks cheaper and is not. Q: Does the ranking hold at fp8? A: fp8 raised throughput by 23 to 31 percent for both vLLM and SGLang on the H100 and lowered cost per token. vLLM stayed ahead. ---