All models

MiMo-V2.5: Xiaomi's multimodal model

1M tokens · Text / Vision / Audio / Video / Code · Prompt cache

Okou no longer runs MiMo-V2.5. This page is kept as a reference for its specs, pricing and benchmarks. For the same kind of work, use GPT 5.6 Luna.

See GPT 5.6 Luna

Use it for cost-sensitive agents that need to inspect mixed media alongside text or code. It is not a dedicated coding provider route like Moonshot or a Z.AI direct route like GLM, so keep Sonnet or Kimi in reserve when tool routing quality is the main constraint.

What is MiMo-V2.5?

June 2026

The model combines a 1M-token context window with text, image, audio, and video input support. That makes it a practical low-cost choice when an agent has to read media-heavy evidence alongside normal text.

Specs at a glance

FamilyXiaomi MiMo
ModalitiesText, image, audio, video, code
LanguagesMultilingual
Context windowUp to 1M tokens
Prompt cachingCache reads supported

MiMo-V2.5 pricing

Provider list price, per 1M tokens.

Input$0.14
Output$0.28
Cache read$0.003
Cache writeNot billed

How MiMo-V2.5 behaves in practice

Observed behaviour from production agent runs.

Multimodal coverage

MiMo-V2.5 is the low-cost route to combine text, images, audio, and video inputs in one agent step.

Large context

The 1M-token context window gives agents enough room for long documents, transcripts, and supporting files without aggressive chunking.

Best agent tasks for MiMo-V2.5

Mixed-media research pass

Use MiMo-V2.5 when the agent needs to inspect screenshots, clips, transcripts, and written notes together before producing a structured brief.

Large-context document review

Load long source material with images or media references and ask for cross-document findings without moving immediately to a premium model.

When to skip MiMo-V2.5

Skip MiMo-V2.5 when you need the strongest Claude-style tool routing, or when the workflow is text-only and a cheaper narrow model is sufficient.

MiMo-V2.5 vs other models

MiMo-V2.5 vs Kimi K2.7 Code

Kimi is the stronger coding-focused Moonshot route. MiMo-V2.5 wins when broad multimodal input and a larger 1M context matter more.

MiMo-V2.5 vs Claude Sonnet 4.6

Sonnet remains the safer premium default for tool routing and hard reasoning. MiMo-V2.5 is the lower-cost multimodal exploration route.

MiMo-V2.5 vs Hy3 Preview

Hy3 Preview is cheaper and text-only with 256K context. MiMo-V2.5 is the better fit when image, audio, video, or a 1M window is needed.

Availability of MiMo-V2.5 on Okou

MiMo-V2.5 has been removed from the Okou lineup, so it can no longer be selected in chat or in a workflow, and it is not available with your own API key. GPT 5.6 Luna covers the same cost-saving tier.