RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started

Model Library

Compare hosted model availability, prices, measured speed, and supported capabilities.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy
  1. Model Library
  2. /
  3. Ornith
  4. /
  5. Ornith 1.5 35B

Ornith 1.5 35B

ornith-ai/Ornith-1.5-35B-A3B
Get API keyView docs

Ornith 1.5 35B is an LLM listed in RunInfra Model APIs. RunInfra serves it as ornith-ai/Ornith-1.5-35B-A3B at $0.10 per 1M input tokens and $0.40 per 1M output tokens. Its context window is 262,144 tokens. The API provides OpenAI-compatible chat completions.

Pricing

USD, pay per token

per 1M input tokens
$0.10
per 1M cached input tokens
$0.01
per 1M output tokens
$0.40

Measured performance

Output speed

310output tokens per second, model only

Time to first token

153milliseconds to first reasoning token

Access

Confirm how your client reaches this model.

Provider
Ornith
API compatibility
OpenAI-compatible chat completions
Accepted input
Text and images
Availability
Available
Data retention
Your prompts are never stored and never used for training. Billing keeps token counts, cost, timing, and a request id. An Idempotency-Key you send holds that response body for 24 hours.

Code examples

Set RUNINFRA_GATEWAY_KEY to your workspace API key before using an example.

Trust and provenance

Verify the company and operating credentials behind this API.

RightNowRunInfra is a sub-product of RightNow Research Lab.
  • SOC 2 Type IIAudited access, logging, and incident response.
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsPricingStartupsBenchmarks

Capacity

Check the limits your workload must fit.

Context window
262,144 tokens
Maximum request size
3.5 MB per request
Maximum generated output
32,768 tokens
Streaming time to first token limit
60 seconds
Maximum request duration
240 seconds

Capabilities

See which request modes the API supports.

Tool calling
Supported
JSON mode
Supported
Streaming
Supported
View Full Spec
Gateway compatibility
OpenAI-compatible chat completions for compatible clients and gateways
OpenRouter Provider Monitor
Provider metadata published as ready for discovery. Not proof of a live OpenRouter listing or callability.
Upstream model
View model
Y CombinatorBacked by Y Combinator.
  • NVIDIA InceptionMember of NVIDIA Inception.
  • Docs
    Research
    News
    Contact
    Backed by
    YCombinator
    AICPA Type II
    SOC 2
    NVIDIA Inception ProgramNVIDIA Inception Program
    Ask AI about RunInfra
    Part of RightNow
    SecurityDPAAUPCookiesTermsPrivacy