RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/Hy3
Model reference

Hy3

This reference covers the open mixture reasoning and agent checkpoint with a quantized companion. It separates text-backed benchmark claims from the image-only appendix and keeps provider pricing distinct.

Hy3 is the vendor reference for tencent/Hy3, as of 2026-08-12. Model scale: 295B total; 21B activated; 3.8B MTP; Context length: 256K, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identitytencent/Hy3
SourceRetrieved 2026-08-12
IdentityVendor descriptionHy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team.
SourceRetrieved 2026-08-12
IdentityLineageFollowing the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data.
SourceRetrieved 2026-08-12
IdentityOpenRouter release dateJul 6, 2026
SourceRetrieved 2026-08-12
ArchitectureModel scale295B total; 21B activated; 3.8B MTP
SourceRetrieved 2026-08-12
ArchitectureLayer configuration80 layers excluding MTP; 1 MTP layer
SourceRetrieved 2026-08-12
ArchitectureExpert routing192 experts, top-8 activated
SourceRetrieved 2026-08-12
ArchitectureAttention64 (GQA, 8 KV heads, head dim 128)
SourceRetrieved 2026-08-12
ArchitectureHidden and intermediate sizeshidden 4096; intermediate 13312
SourceRetrieved 2026-08-12
ArchitectureVocabulary120832
SourceRetrieved 2026-08-12
ArchitecturePrecisionBF16
SourceRetrieved 2026-08-12
ArchitectureVendor architecture framingBuilt on a hybrid fast-and-slow-thinking Mixture-of-Experts (MoE) architecture
SourceRetrieved 2026-08-12
ContextContext length256K
SourceRetrieved 2026-08-12
ModalitiesPipeline tagtext generation
SourceRetrieved 2026-08-12
LicenseWeights licenseHy3 is released under the Apache License 2.0.
SourceRetrieved 2026-08-12
LicenseLicense fileApache License, Version 2.0; Copyright (C) 2026 Tencent. All rights reserved.
SourceRetrieved 2026-08-12
PricingOpenRouter listing$0.1288 per 1M input tokens / $0.5336 per 1M output tokens
SourceRetrieved 2026-08-12
AvailabilityRepository accessUngated repository.
SourceRetrieved 2026-08-12
AvailabilityWeight distributionWe open-source Hy3 and Hy3-FP8 model weights on Hugging Face, ModelScope, GitCode, and CNB.
SourceRetrieved 2026-08-12
AvailabilityFP8 companiontencent/Hy3-FP8; FP8 quantized instruct model
SourceRetrieved 2026-08-12
AvailabilityWorkBuddy accessavailable free of charge to users worldwide until 31 August 2026 (Pacific Time)
SourceRetrieved 2026-08-12
AvailabilityvLLM recipe floorvLLM 0.26.0+
SourceRetrieved 2026-08-12
AvailabilityVendor serving recommendationFor production serving, we recommend using vLLM or SGLang, both of which provide dedicated recipes for Hy3
SourceRetrieved 2026-08-12
AvailabilityMain weightstencent/Hy3
SourceRetrieved 2026-08-12
AvailabilityQuantized instruct weightstencent/Hy3-FP8
SourceRetrieved 2026-08-12

What the vendor says is new

"Today, we introduce Hy3, which outperforms similar-size models and rivals flagship open-source models with 2-5x parameters."

SourceRetrieved 2026-08-12

"Hy3 recorded more than 68 times as many API calls as the previous-generation model and ranked first globally on OpenRouter's global LLM usage leaderboard within one week of launch."

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

BenchmarkVendor-claimed valueSource and date
Expert blind evaluationHy3 scored 2.67/4, outperforming GLM-5.1 at 2.51/4
SourceRetrieved 2026-08-12
Internal hallucination evaluationdropped from 12.5% to 5.4% in internal evaluations
SourceRetrieved 2026-08-12

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

vLLM

The [Model] Support Hy3 preview pull request was merged on April 23, 2026.

SourceRetrieved 2026-08-12

SGLang

The Hy3 day-zero cookbook pull request was merged on July 6, 2026.

SourceRetrieved 2026-08-12

Serving concepts

  • Mixture-of-experts serving->
  • Expert parallelism->
  • Grouped-query attention->
  • Context length->
  • FP8 quantization->
  • KV cache->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is Hy3?

As of 2026-08-12, release identity: tencent/Hy3.

Where are the Hy3 weights?

As of 2026-08-12, repository access: Ungated repository.

Can I run Hy3 myself?

As of 2026-08-12, repository access: Ungated repository.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy