Skip to content

MiniCPM5-2B

POST /v1/chat/completions model: MiniCPM5-2B

The weights served here are openbmb/MiniCPM5-2B, read from the config.json shipped with them:

ArchitectureLlamaForCausalLM, model_type = llama
Layers42
Hidden size2,048
Attention16 query heads, 2 KV heads (GQA)
Intermediate size6,144
Vocabulary130,560
Context131,072
Precisionbfloat16 — not quantised

The only dense (non-MoE) model on this endpoint, and the only unquantised one. Its 131,072-token context is the shortest here.

Context length131,072 tokens
Input modalitiestext only
Streaming✅
Tool calling✅ (see below)
JSON output✅ json_object
Thinking❌
EnginevLLM
Stabilityexperimental

This model answers directly. It does not return separated thinking, so choices[0].message.reasoning stays empty and usage.completion_tokens_details.reasoning_tokens is 0.

Read the answer from choices[0].message.content. For separated thinking, use DeepSeek-V4-Flash or Qwen3.8-Flash-Next.

A system message may sit at any position, and there may be more than one. Spell the role system.

Both tools and parallel_tool_calls are accepted. Offered a get_weather tool and asked for the weather in Paris, the model issued the call and returned finish_reason: "tool_calls". As with any model this size, try it against your own prompts before relying on it for tool-driven work.

Context window131,072 tokens, counted as prompt plus output
JSON outputresponse_format: {"type": "json_object"}
Inputtext only