MiMo-V2.6-Flash
MiMo-V2.6-Flash
About the model
Section titled “About the model”The weights served here are MiMo-V2.6-Flash, the efficiency-balanced member of Xiaomi’s MiMo-V2.6 line — a sparse Mixture-of-Experts model that accepts text, images, video and audio in one endpoint, and thinks before answering unless told not to.
Architecture
Section titled “Architecture”From the shipped config.json:
| Architecture | MiMoV2ForCausalLM, model_type = mimo_v2 |
| Parameters | 309B total, 15B activated per token |
| Layers | 48 — 39 sliding-window + 9 global; the first block is global attention with a dense FFN, the other 47 are MoE |
| Hidden dimension | 4,096 |
| Attention | 64 query heads; 4 KV heads on the global layers, 8 on the sliding-window ones |
| Head dimensions | 192 for Q/K, 128 for V — asymmetric |
| Sliding window | 128 tokens |
| Mixture of Experts | 256 routed experts, 8 activated per token, expert intermediate 2,048, no shared expert |
| Multi-token prediction | 3 layers declared in config.json; the shipped dflash/ drafter is 5 layers |
| Vocabulary | 152,576 |
| Native context | 1,048,576 |
| Quantisation | quant_method: fp8 e4m3, block [128, 128], stored as MXFP4 (store_dtype: mxfp4, block 32) |
| Licence | MIT |
The vision tower is a 681M-parameter MiMo ViT (28 layers, 24 sliding-window + 4 full, patch 16, spatial merge 2×2). Audio goes through a 308M AudioTokenizer plus a 127M patch encoder.
On this endpoint
Section titled “On this endpoint”At a glance
Section titled “At a glance”| Context length | 1,048,576 tokens |
| Input modalities | text, image, audio |
| Streaming | ✅ |
| Tool calling | ✅ |
| JSON output | ✅ json_object and json_schema |
| Thinking | ✅ — on unless you turn it off |
| Inference engine | SGLang |
| Stability | experimental |
This is the only model on this API that accepts audio, and the only one that accepts a
strict json_schema rather than just json_object.
messages
Section titled “messages”A system message may sit at any position, and there may be more than one.
Thinking
Section titled “Thinking”Thinking is on by default. Omit reasoning_effort entirely and the response still carries a
populated reasoning:
"message": { "role": "assistant", "content": "4", "reasoning": "2+2 is 4."}To turn it off, send reasoning_effort: "none" — reasoning then comes back empty.
Image input
Section titled “Image input”Images are passed as image_url parts, base64 data URLs included:
{ "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "data:image/png;base64,..." } }, { "type": "text", "text": "How many blue circles?" } ]}usage.prompt_tokens_details.image_tokens reports what the image cost — a 1200×420 PNG came
back as 494 image tokens.
Audio input
Section titled “Audio input”Audio uses the OpenAI input_audio part:
{ "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": "<base64>", "format": "wav" } }, { "type": "text", "text": "What is this sound?" } ]}The audio front end resamples to 24 kHz. No other model on this API accepts this part type.
Structured output
Section titled “Structured output”Both forms work:
response_format | Result |
|---|---|
{"type": "json_object"} | ✅ {"北京": 2174, "上海": 2487} |
{"type": "json_schema", "json_schema": {..., "strict": true}} | ✅ {"city":"北京","pop":2174} |
Limits
Section titled “Limits”| Context window | 1,048,576 tokens, counted as a total budget — prompt plus output |
| Input | text, image, audio |
| JSON output | json_object and strict json_schema |
| Thinking tiers | reasoning_effort — send none to disable; other tiers are accepted but the model is not tier-calibrated the way the DeepSeek models are |
Example
Section titled “Example”curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \ -H "Authorization: Bearer $RADEON_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "MiMo-V2.6-Flash", "messages": [ { "role": "user", "content": "Why is the sky blue?" } ] }'Add "reasoning_effort": "none" to get the answer without the thinking trace.