MiMo-V2.5: Xiaomi's multimodal model
1M tokens · Text / Vision / Audio / Video / Code · Prompt cache
Okou no longer runs MiMo-V2.5. This page is kept as a reference for its specs, pricing and benchmarks. For the same kind of work, use GPT 5.6 Luna.
See GPT 5.6 LunaUse it for cost-sensitive agents that need to inspect mixed media alongside text or code. It is not a dedicated coding provider route like Moonshot or a Z.AI direct route like GLM, so keep Sonnet or Kimi in reserve when tool routing quality is the main constraint.
What is MiMo-V2.5?
June 2026
The model combines a 1M-token context window with text, image, audio, and video input support. That makes it a practical low-cost choice when an agent has to read media-heavy evidence alongside normal text.
Specs at a glance
MiMo-V2.5 pricing
Provider list price, per 1M tokens.
How MiMo-V2.5 behaves in practice
Observed behaviour from production agent runs.
Multimodal coverage
MiMo-V2.5 is the low-cost route to combine text, images, audio, and video inputs in one agent step.
Large context
The 1M-token context window gives agents enough room for long documents, transcripts, and supporting files without aggressive chunking.
Best agent tasks for MiMo-V2.5
Mixed-media research pass
Use MiMo-V2.5 when the agent needs to inspect screenshots, clips, transcripts, and written notes together before producing a structured brief.
Large-context document review
Load long source material with images or media references and ask for cross-document findings without moving immediately to a premium model.
When to skip MiMo-V2.5
Skip MiMo-V2.5 when you need the strongest Claude-style tool routing, or when the workflow is text-only and a cheaper narrow model is sufficient.
MiMo-V2.5 vs other models
MiMo-V2.5 vs Kimi K2.7 Code
Kimi is the stronger coding-focused Moonshot route. MiMo-V2.5 wins when broad multimodal input and a larger 1M context matter more.
MiMo-V2.5 vs Claude Sonnet 4.6
Sonnet remains the safer premium default for tool routing and hard reasoning. MiMo-V2.5 is the lower-cost multimodal exploration route.
MiMo-V2.5 vs Hy3 Preview
Hy3 Preview is cheaper and text-only with 256K context. MiMo-V2.5 is the better fit when image, audio, video, or a 1M window is needed.
Availability of MiMo-V2.5 on Okou
MiMo-V2.5 has been removed from the Okou lineup, so it can no longer be selected in chat or in a workflow, and it is not available with your own API key. GPT 5.6 Luna covers the same cost-saving tier.