Skip to content

MiMo-V2.6-Flash

POST /v1/chat/completions model: MiMo-V2.6-Flash

The weights served here are MiMo-V2.6-Flash, the efficiency-balanced member of Xiaomi’s MiMo-V2.6 line — a sparse Mixture-of-Experts model that accepts text, images, video and audio in one endpoint, and thinks before answering unless told not to.

From the shipped config.json:

ArchitectureMiMoV2ForCausalLM, model_type = mimo_v2
Parameters309B total, 15B activated per token
Layers48 — 39 sliding-window + 9 global; the first block is global attention with a dense FFN, the other 47 are MoE
Hidden dimension4,096
Attention64 query heads; 4 KV heads on the global layers, 8 on the sliding-window ones
Head dimensions192 for Q/K, 128 for V — asymmetric
Sliding window128 tokens
Mixture of Experts256 routed experts, 8 activated per token, expert intermediate 2,048, no shared expert
Multi-token prediction3 layers declared in config.json; the shipped dflash/ drafter is 5 layers
Vocabulary152,576
Native context1,048,576
Quantisationquant_method: fp8 e4m3, block [128, 128], stored as MXFP4 (store_dtype: mxfp4, block 32)
LicenceMIT

The vision tower is a 681M-parameter MiMo ViT (28 layers, 24 sliding-window + 4 full, patch 16, spatial merge 2×2). Audio goes through a 308M AudioTokenizer plus a 127M patch encoder.

Context length1,048,576 tokens
Input modalitiestext, image, audio
Streaming✅
Tool calling✅
JSON output✅ json_object and json_schema
Thinking✅ — on unless you turn it off
Inference engineSGLang
Stabilityexperimental

This is the only model on this API that accepts audio, and the only one that accepts a strict json_schema rather than just json_object.

A system message may sit at any position, and there may be more than one.

Thinking is on by default. Omit reasoning_effort entirely and the response still carries a populated reasoning:

"message": {
"role": "assistant",
"content": "4",
"reasoning": "2+2 is 4."
}

To turn it off, send reasoning_effort: "none" — reasoning then comes back empty.

Images are passed as image_url parts, base64 data URLs included:

{
"role": "user",
"content": [
{ "type": "image_url", "image_url": { "url": "data:image/png;base64,..." } },
{ "type": "text", "text": "How many blue circles?" }
]
}

usage.prompt_tokens_details.image_tokens reports what the image cost — a 1200×420 PNG came back as 494 image tokens.

Audio uses the OpenAI input_audio part:

{
"role": "user",
"content": [
{ "type": "input_audio", "input_audio": { "data": "<base64>", "format": "wav" } },
{ "type": "text", "text": "What is this sound?" }
]
}

The audio front end resamples to 24 kHz. No other model on this API accepts this part type.

Both forms work:

response_formatResult
{"type": "json_object"}✅ {"北京": 2174, "上海": 2487}
{"type": "json_schema", "json_schema": {..., "strict": true}}✅ {"city":"北京","pop":2174}
Context window1,048,576 tokens, counted as a total budget — prompt plus output
Inputtext, image, audio
JSON outputjson_object and strict json_schema
Thinking tiersreasoning_effort — send none to disable; other tiers are accepted but the model is not tier-calibrated the way the DeepSeek models are
Terminal window
curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \
-H "Authorization: Bearer $RADEON_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MiMo-V2.6-Flash",
"messages": [
{ "role": "user", "content": "Why is the sky blue?" }
]
}'

Add "reasoning_effort": "none" to get the answer without the thinking trace.