RunInfraby RightNow
  • Model APIsNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Loading related term
Loading related term
Loading related term
Loading related term
Home/Catalog/MiMo-V2.5
Model reference

MiMo-V2.5

This reference covers the open omnimodal mixture checkpoint and keeps sibling scale claims separate. It records the vendor's permissive license statement without implying that a repository license file exists.

MiMo-V2.5 is the vendor reference for XiaomiMiMo/MiMo-V2.5, as of 2026-08-12. Model scale: Sparse MoE (Mixture of Experts), 310B total / 15B activated parameters; Context length: Up to 1M tokens, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.

Cited specifications

GroupFactSource-cited display valueSource and date
IdentityRelease identityXiaomiMiMo/MiMo-V2.5
SourceRetrieved 2026-08-12
IdentityVendor descriptionMiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture.
SourceRetrieved 2026-08-12
IdentityOpenRouter release dateApr 22, 2026
SourceRetrieved 2026-08-12
IdentityOpen-source announcementToday, we officially open source the Xiaomi MiMo-V2.5 series
SourceRetrieved 2026-08-12
IdentityPublic testing datePublic testing began April 23, 2026.
SourceRetrieved 2026-08-12
ArchitectureModel scaleSparse MoE (Mixture of Experts), 310B total / 15B activated parameters
SourceRetrieved 2026-08-12
ArchitectureExpert routing256 routed experts; 8 per token
SourceRetrieved 2026-08-12
ArchitectureLayer configuration48 layers (1 dense + 47 MoE)
SourceRetrieved 2026-08-12
ArchitectureAttention layers9 full-attention + 39 SWA
SourceRetrieved 2026-08-12
ArchitectureAttention interleavinginterleaving Sliding Window Attention (SWA) and Global Attention (GA) with a 5:1 ratio and 128 sliding window. This reduces KV-cache storage by nearly 6x while maintaining long-context performance via learnable attention sink bias.
SourceRetrieved 2026-08-12
ArchitectureVision encoder729M-param ViT (28 layers: 24 SWA + 4 Full)
SourceRetrieved 2026-08-12
ArchitectureAudio encoder261M-param Audio Transformer
SourceRetrieved 2026-08-12
ArchitectureMulti-token prediction329M parameters, 3 layers, for speculative decoding
SourceRetrieved 2026-08-12
ArchitectureTrainingTrained on a total of ~48T tokens using FP8 mixed precision.
SourceRetrieved 2026-08-12
ArchitectureBackbone lineageinherits from the MiMo-V2-Flash architecture
SourceRetrieved 2026-08-12
ContextContext lengthUp to 1M tokens
SourceRetrieved 2026-08-12
ContextContext extension schedule32K -> 256K -> 1M
SourceRetrieved 2026-08-12
ModalitiesInput understandingtext, image, video, and audio
SourceRetrieved 2026-08-12
ModalitiesOutputtext
SourceRetrieved 2026-08-12
LicenseVendor license statementuses the MIT license, supports commercial inference deployment and secondary training, and requires no additional authorization.
SourceRetrieved 2026-08-12
LicenseRepository metadatalicense: mit
SourceRetrieved 2026-08-12
PricingXiaomi cache-hit input$0.0028 / MTok
SourceRetrieved 2026-08-12
PricingXiaomi cache-miss input$0.14 / MTok
SourceRetrieved 2026-08-12
PricingXiaomi output$0.28 / MTok
SourceRetrieved 2026-08-12
PricingOpenRouter listing$0.14 input and $0.28 output per 1M tokens
SourceRetrieved 2026-08-12
AvailabilityRepository accessUngated repository.
SourceRetrieved 2026-08-12
AvailabilityTensor typesF32, BF16, F8_E4M3
SourceRetrieved 2026-08-12
AvailabilityConfiguration refreshThe config.json and tokenizer_config.json files in this repository have been updated since the initial release... Using the outdated config may lead to degraded model performance.
SourceRetrieved 2026-08-12
AvailabilityRecommended samplingtemperature=1.0, top_p=0.95
SourceRetrieved 2026-08-12
AvailabilityBase siblingMiMo-V2.5-Base; 256K context
SourceRetrieved 2026-08-12
AvailabilityPro siblingMiMo-V2.5-Pro; 1.02T total / 42B activated
SourceRetrieved 2026-08-12
AvailabilityvLLM recipe floorvLLM 0.21.0+
SourceRetrieved 2026-08-12
AvailabilityMain weightsXiaomiMiMo/MiMo-V2.5
SourceRetrieved 2026-08-12

What the vendor says is new

"Today, we are releasing MiMo-V2.5, a major step forward in agentic capability and multimodal understanding. With native visual and audio understanding, MiMo-V2.5 reasons seamlessly across modalities, surpasses MiMo-V2-Pro in agentic performance, and supports up to 1 million tokens of context."

SourceRetrieved 2026-08-12

"Post-training incorporates SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD), achieving strong performance on agentic tasks and multimodal understanding benchmarks."

SourceRetrieved 2026-08-12

Vendor-claimed benchmarks

These results are vendor-claimed, not independently measured by RunInfra.

BenchmarkVendor-claimed valueSource and date
Claw-Eval, general subset62.3; placing it at the Pareto frontier of performance and efficiency
SourceRetrieved 2026-08-12

Serving support

Listed rows have a cited upstream support signal. They are not RunInfra measurements.

SGLang

The [Feature] Xiaomi MiMo-V2.5 day0 support pull request was merged on April 30, 2026.

SourceRetrieved 2026-08-12

vLLM

The [Model] Add MiMo-V2.5 support pull request was merged on April 27, 2026.

SourceRetrieved 2026-08-12

Serving concepts

  • Mixture-of-experts serving->
  • Sliding-window attention->
  • Speculative decoding->
  • FP8 quantization->
  • Context length->
  • KV cache->

What RunInfra measures when we measure it

RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.

  • Measurement methodology->
  • Measured package catalog->
  • Published benchmarks->

Questions about this model reference

What is MiMo-V2.5?

As of 2026-08-12, release identity: XiaomiMiMo/MiMo-V2.5.

Where are the MiMo-V2.5 weights?

As of 2026-08-12, repository access: Ungated repository.

Can I run MiMo-V2.5 myself?

As of 2026-08-12, repository access: Ungated repository.

RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy