SGLang
The [Feature] Xiaomi MiMo-V2.5 day0 support pull request was merged on April 30, 2026.
This reference covers the open omnimodal mixture checkpoint and keeps sibling scale claims separate. It records the vendor's permissive license statement without implying that a repository license file exists.
MiMo-V2.5 is the vendor reference for XiaomiMiMo/MiMo-V2.5, as of 2026-08-12. Model scale: Sparse MoE (Mixture of Experts), 310B total / 15B activated parameters; Context length: Up to 1M tokens, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | XiaomiMiMo/MiMo-V2.5 | SourceRetrieved 2026-08-12 |
| Identity | Vendor description | MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. | SourceRetrieved 2026-08-12 |
| Identity | OpenRouter release date | Apr 22, 2026 | SourceRetrieved 2026-08-12 |
| Identity | Open-source announcement | Today, we officially open source the Xiaomi MiMo-V2.5 series | SourceRetrieved 2026-08-12 |
| Identity | Public testing date | Public testing began April 23, 2026. | SourceRetrieved 2026-08-12 |
| Architecture | Model scale | Sparse MoE (Mixture of Experts), 310B total / 15B activated parameters | SourceRetrieved 2026-08-12 |
| Architecture | Expert routing | 256 routed experts; 8 per token | SourceRetrieved 2026-08-12 |
| Architecture | Layer configuration | 48 layers (1 dense + 47 MoE) | SourceRetrieved 2026-08-12 |
| Architecture | Attention layers | 9 full-attention + 39 SWA | SourceRetrieved 2026-08-12 |
| Architecture | Attention interleaving | interleaving Sliding Window Attention (SWA) and Global Attention (GA) with a 5:1 ratio and 128 sliding window. This reduces KV-cache storage by nearly 6x while maintaining long-context performance via learnable attention sink bias. | SourceRetrieved 2026-08-12 |
| Architecture | Vision encoder | 729M-param ViT (28 layers: 24 SWA + 4 Full) | SourceRetrieved 2026-08-12 |
| Architecture | Audio encoder | 261M-param Audio Transformer | SourceRetrieved 2026-08-12 |
| Architecture | Multi-token prediction | 329M parameters, 3 layers, for speculative decoding | SourceRetrieved 2026-08-12 |
| Architecture | Training | Trained on a total of ~48T tokens using FP8 mixed precision. | SourceRetrieved 2026-08-12 |
| Architecture | Backbone lineage | inherits from the MiMo-V2-Flash architecture | SourceRetrieved 2026-08-12 |
| Context | Context length | Up to 1M tokens | SourceRetrieved 2026-08-12 |
| Context | Context extension schedule | 32K -> 256K -> 1M | SourceRetrieved 2026-08-12 |
| Modalities | Input understanding | text, image, video, and audio | SourceRetrieved 2026-08-12 |
| Modalities | Output | text | SourceRetrieved 2026-08-12 |
| License | Vendor license statement | uses the MIT license, supports commercial inference deployment and secondary training, and requires no additional authorization. | SourceRetrieved 2026-08-12 |
| License | Repository metadata | license: mit | SourceRetrieved 2026-08-12 |
| Pricing | Xiaomi cache-hit input | $0.0028 / MTok | SourceRetrieved 2026-08-12 |
| Pricing | Xiaomi cache-miss input | $0.14 / MTok | SourceRetrieved 2026-08-12 |
| Pricing | Xiaomi output | $0.28 / MTok | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter listing | $0.14 input and $0.28 output per 1M tokens | SourceRetrieved 2026-08-12 |
| Availability | Repository access | Ungated repository. | SourceRetrieved 2026-08-12 |
| Availability | Tensor types | F32, BF16, F8_E4M3 | SourceRetrieved 2026-08-12 |
| Availability | Configuration refresh | The config.json and tokenizer_config.json files in this repository have been updated since the initial release... Using the outdated config may lead to degraded model performance. | SourceRetrieved 2026-08-12 |
| Availability | Recommended sampling | temperature=1.0, top_p=0.95 | SourceRetrieved 2026-08-12 |
| Availability | Base sibling | MiMo-V2.5-Base; 256K context | SourceRetrieved 2026-08-12 |
| Availability | Pro sibling | MiMo-V2.5-Pro; 1.02T total / 42B activated | SourceRetrieved 2026-08-12 |
| Availability | vLLM recipe floor | vLLM 0.21.0+ | SourceRetrieved 2026-08-12 |
| Availability | Main weights | XiaomiMiMo/MiMo-V2.5 | SourceRetrieved 2026-08-12 |
"Today, we are releasing MiMo-V2.5, a major step forward in agentic capability and multimodal understanding. With native visual and audio understanding, MiMo-V2.5 reasons seamlessly across modalities, surpasses MiMo-V2-Pro in agentic performance, and supports up to 1 million tokens of context."
SourceRetrieved 2026-08-12
"Post-training incorporates SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD), achieving strong performance on agentic tasks and multimodal understanding benchmarks."
SourceRetrieved 2026-08-12
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| Claw-Eval, general subset | 62.3; placing it at the Pareto frontier of performance and efficiency | SourceRetrieved 2026-08-12 |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: XiaomiMiMo/MiMo-V2.5.
As of 2026-08-12, repository access: Ungated repository.
As of 2026-08-12, repository access: Ungated repository.
© 2026 RunInfra. All rights reserved.