RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Model Library

Compare hosted model availability, prices, measured speed, and supported capabilities.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
  1. Model Library
  2. /
  3. Qwen
  4. /
  5. Qwen3.8 27B

Qwen3.8 27B

Qwen/Qwen3.8-27B
Get API keyView docs

Qwen3.8 27B is an LLM listed in RunInfra Model APIs. RunInfra serves it as Qwen/Qwen3.8-27B at $0.00 per 1M input and output tokens until Aug 18, 2026, 11:00 AM UTC, with standard rates resuming automatically. Its context window is 262,144 tokens. The API provides OpenAI-compatible chat completions.

Pricing

USD, pay per token

per 1M input and output tokens
$0.00

Input and output tokens are listed at $0.00 until Aug 18, 2026, 11:00 AM UTC.

02days
20hrs
03min
55sec
Ends in 2 days 20 hours

Standard rates resume automatically: $0.10 per 1M input tokens, $0.40 per 1M output tokens.

Hosted inference is available to every account. You pay for token usage from the same workspace balance.

Measured performance

Output speed

140output tokens per second, end to end

Time to first token

939milliseconds to first reasoning token

Access

Confirm how your client reaches this model.

Provider
Qwen
API compatibility
OpenAI-compatible chat completions
Availability
Available

Capacity

Check the limits your workload must fit.

Context window
262,144 tokens
Maximum request size
3.5 MB per request
Maximum generated output
32,768 tokens
Streaming time to first token limit
60 seconds
Maximum request duration
240 seconds

Capabilities

See which request modes the API supports.

Tool calling
Supported
JSON mode
Supported
Streaming
Supported
View Full Spec
Gateway compatibility
OpenAI-compatible chat completions for compatible clients and gateways
OpenRouter Provider Monitor
Provider metadata published as ready for discovery. Not proof of a live OpenRouter listing or callability.
Upstream model
View model

Code examples

Set RUNINFRA_GATEWAY_KEY to your workspace API key before using an example.

Trust and provenance

Verify the company and operating credentials behind this API.

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
  • Y CombinatorBacked by Y Combinator.
  • NVIDIA InceptionMember of NVIDIA Inception.
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy