MiniCPM5-2B
MiniCPM5-2B
Specification
Section titled “Specification”The weights served here are openbmb/MiniCPM5-2B, read from the config.json shipped with
them:
| Architecture | LlamaForCausalLM, model_type = llama |
| Layers | 42 |
| Hidden size | 2,048 |
| Attention | 16 query heads, 2 KV heads (GQA) |
| Intermediate size | 6,144 |
| Vocabulary | 130,560 |
| Context | 131,072 |
| Precision | bfloat16 — not quantised |
The only dense (non-MoE) model on this endpoint, and the only unquantised one. Its 131,072-token context is the shortest here.
On this endpoint
Section titled “On this endpoint”At a glance
Section titled “At a glance”| Context length | 131,072 tokens |
| Input modalities | text only |
| Streaming | ✅ |
| Tool calling | ✅ (see below) |
| JSON output | ✅ json_object |
| Thinking | ❌ |
| Engine | vLLM |
| Stability | experimental |
Thinking
Section titled “Thinking”This model answers directly. It does not return separated thinking, so
choices[0].message.reasoning stays empty and
usage.completion_tokens_details.reasoning_tokens is 0.
Read the answer from choices[0].message.content. For separated thinking, use
DeepSeek-V4-Flash or
Qwen3.8-Flash-Next.
messages
Section titled “messages”A system message may sit at any position, and there may be more than one. Spell the role
system.
Both tools and parallel_tool_calls are accepted. Offered a get_weather tool and asked for
the weather in Paris, the model issued the call and returned
finish_reason: "tool_calls". As with any model this size, try it against your own prompts
before relying on it for tool-driven work.
Limits
Section titled “Limits”| Context window | 131,072 tokens, counted as prompt plus output |
| JSON output | response_format: {"type": "json_object"} |
| Input | text only |