MiMo hands-on: the token savings are real, but what else?
Community Discussion · Tracks

MiMo hands-on: the token savings are real, but what else?

Sister Liang on ValuationSister Liang on ValuationSep 252026/09/25 147 views

As someone who knows nothing about Xiaomi's model ecosystem, I gave MiMo a try. Let me be upfront: this isn't my first time using large models. I've used OpenAI for about two months, DeepSeek Harness for roughly a month, and I've also dabbled with on-device models for a month. But Xiaomi's MiMo is genuinely my first time, so this post is just about what I ran into and what I saw—I'm not shilling for anyone.

There are two entry points: a web version for direct chat, and an API integration. I went through both. The web version has no barrier—you go in and there's just an input box, looks like most chat products, no configuration needed. The API route takes a bit more effort: you have to add credentials first, then confirm whether the model ID shows up in the current config. I got stuck here for a moment—the model catalog updates lag behind the release, so you have to check yourself whether the target ID is listed.

MiMo V2.5 Pro has one great thing going for it: it seems to be the most token-efficient open-source model right now.

That's from an InfoQ review. I ran a few long-text tasks with it, and my impression matches. It doesn't talk much when thinking, doesn't beat around the bush. For the same chunk of material, its answer comes out shorter than the ones I usually use, without missing key points.

First, long context. The V2.5 series claims to support 1 million tokens of context. I can't feed that much in one go, but I threw in a long report plus a few attachments, and it didn't crash or start making things up. Anyone doing due diligence knows what that means—prospectus plus research reports, can be done in one pass.

Now token efficiency. According to ClawEval's tests, V2.5-Pro consumes about 70k tokens per trajectory, achieving 64% Pass^3, using roughly 40-60% fewer tokens than Claude Opus 4.6, Gemini 3.1 Pro, and GPT-5.4. I can't reproduce that number, but my own bookkeeping feels consistent—running the same tasks, the bill is indeed lower.

Open source is another point. MIT license, 1T total parameters, 42B activated parameters—this combo means teams wanting to self-deploy have options, not locked into API pricing. It also has multimodal versions for text, vision, and voice, following a different context spec. MiMo isn't just a cloud model either—it's meant to be embedded in phones, Xiao Ai, and car cockpits. I can't see the internals, but logically, once on-device inference works, the cost structure is a different beast from pure cloud.

There are downsides too. Coding isn't its strongest suit. Some evals ran frontend logic, full-stack apps, Java backend—conclusion is it's capable, but not top-tier. I wrote a few scripts with it, usable, but for complex refactoring I'd switch back to my usual tools. Documentation and ecosystem are still catching up—small things like the model ID catalog lagging will trip up first-time integrators. The maturity of toolchains, SDKs, and peripheral integrations isn't on the same level as OpenAI. The version cadence is also fast—V2.5, V2.6, Pro, Flash—naming and versions stack up densely, easy for people doing selection to get dizzy.

So my judgment is: it depends. Suitable for teams already in the Xiaomi ecosystem, for budget-sensitive people running long-horizon agent tasks and sensitive to token costs, and for open-source advocates wanting to self-deploy. Not suitable for those who make coding ability the sole criterion, nor for those whose toolchains are already locked into another ecosystem where migration cost exceeds benefit.

On trends: the next watershed for open-source models is token efficiency, not parameters. Whoever's model can halve the cost on a task will get into enterprise procurement lists. The valuation logic of model companies will shift from "how big" to "how much per task," and the on-device line will be the first to settle this account.

2 replies

?
Ctrl + Enter to reply
Lin
LinSep 26

I've also gotten stuck on model ID directory lag. The very first hurdle of hooking up the API is enough to scare people off.

hongtao
hongtaoSep 26

Is this direction even good for publishing papers... a 42B activated-parameter deployment cost is something a lab can't afford. Funding is tight.