RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Model Library

Compare hosted model availability, prices, measured speed, and supported capabilities.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
  1. Model Library
  2. /
  3. Qwen
  4. /
  5. Qwen3.8 2.4T A95B

Qwen3.8 2.4T A95B

Inferact/Qwen3.8-2.4T-A95B-NVFP4
Get API keyView docs

Qwen3.8 2.4T A95B is an LLM listed in RunInfra Model APIs. RunInfra serves it as Inferact/Qwen3.8-2.4T-A95B-NVFP4 at $2.00 per 1M input tokens and $6.00 per 1M output tokens. Its context window is 262,144 tokens. The API provides OpenAI-compatible chat completions.

Pricing

USD, pay per token

per 1M input tokens
$2.00
per 1M output tokens
$6.00

Hosted inference is available to every account. You pay for token usage from the same workspace balance.

Access

Confirm how your client reaches this model.

Provider
Qwen
API compatibility
OpenAI-compatible chat completions
Availability
Available

Capacity

Check the limits your workload must fit.

Context window
262,144 tokens
Maximum request size
3.5 MB per request
Maximum generated output
32,768 tokens
Streaming time to first token limit
60 seconds
Maximum request duration
300 seconds

Capabilities

See which request modes the API supports.

Tool calling
Supported
Streaming
Supported
Served precision
NVFP4, routed experts only
View Full Spec
Gateway compatibility
OpenAI-compatible chat completions for compatible clients and gateways
OpenRouter Provider Monitor
Provider metadata published as ready for discovery. Not proof of a live OpenRouter listing or callability.
Upstream model
View model

Code examples

Set RUNINFRA_GATEWAY_KEY to your workspace API key before using an example.

Trust and provenance

Verify the company and operating credentials behind this API.

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
  • Y CombinatorBacked by Y Combinator.
  • NVIDIA InceptionMember of NVIDIA Inception.
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy