GLM-5.3-Flash
GLM-5.3-Flash
About the model
Section titled “About the model”The weights served here are GLM-5.3-Flash, Z.AI’s speed-oriented member of the GLM-5.3 line — a sparse Mixture-of-Experts model that thinks before answering, and does so without being asked.
Architecture
Section titled “Architecture”From the shipped config.json:
| Architecture | Glm5NextForConditionalGeneration, model_type = glm5_next |
| Layers | 45 — the first 3 dense, the rest MoE |
| Hidden dimension | 4,096 |
| Attention | 64 query heads, 64 KV heads — full MHA, not grouped |
| Dense intermediate | 12,288 |
| Mixture of Experts | 288 routed experts + 1 shared, 8 routed activated per token, expert intermediate 2,048 |
| Multi-token prediction | 1 layer |
| Vocabulary | 154,880 |
| Native context | 1,048,576 |
| Quantisation | FP8, e4m3, block [128, 128], dynamic activation scaling |
| Licence | MIT |
On this endpoint
Section titled “On this endpoint”At a glance
Section titled “At a glance”| Context length | 262,144 tokens |
| Input modalities | text only |
| Streaming | ✅ |
| Tool calling | ✅ |
| JSON output | ✅ json_object |
| Thinking | ✅ — on unless you turn it down |
| Inference engine | vLLM |
| Stability | experimental |
messages
Section titled “messages”A system message may sit at any position, and there may be more than one — unlike
Qwen3.8-Flash-Next, which allows exactly one and
insists it comes first.
Thinking
Section titled “Thinking”Omitting reasoning_effort does not disable thinking. A plain request already returns a
populated reasoning. How long the model thinks tracks the prompt rather than the tier — a
one-line question may produce a few dozen reasoning tokens, a puzzle several thousand.
Accepted values:
| Tier | Notes |
|---|---|
| omitted | still thinks |
low | |
medium | |
high | the longest tier |
There is no reasoning_effort value that turns thinking off — none is not in the enum. To keep
a response short, cap max_tokens instead, and read the answer from content rather than
reasoning.
Where the output lands:
| Thinking text | choices[0].message.reasoning |
| Token count | usage.completion_tokens_details.reasoning_tokens |
Image input
Section titled “Image input”Not available on this endpoint, despite the vision tower in the weights. An image_url part comes
back as a 400 whose message describes a server-side path setting rather than the real reason:
Invalid `--allowed-local-media-path`: The path <path> does not exist.Send text only. For images use DeepSeek-V4-Flash-Vision-Exp, DeepSeek-V4.1-Flash or either Qwen3.8 model.
Limits
Section titled “Limits”| Context window | 262,144 tokens, counted as a total budget — prompt plus output, not an output-only allowance |
| Input | text only; for images use DeepSeek-V4.1-Flash |
| JSON output | response_format: {"type": "json_object"} |
| Thinking tiers | reasoning_effort (or the equivalent reasoning.effort) — low, medium, high |
max_tokens is capped against the total budget, prompt included:
This model's maximum context length is 262144 tokens. However, you requested 128000 output tokens and your prompt contains at least 134145 input tokens, for a total of at least 262145 tokens. Please reduce the length of the input prompt or the number of requested output tokens.That one is a 400, unlike the 422 above.
Example
Section titled “Example”curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \ -H "Authorization: Bearer $RADEON_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "GLM-5.3-Flash", "reasoning_effort": "high", "messages": [ { "role": "user", "content": "Why is the sky blue?" } ] }'Swap high for xhigh and the same request becomes a 422 on this model — see
Thinking.