Skip to content

GLM-5.3-Flash

POST /v1/chat/completions model: GLM-5.3-Flash

The weights served here are GLM-5.3-Flash, Z.AI’s speed-oriented member of the GLM-5.3 line — a sparse Mixture-of-Experts model that thinks before answering, and does so without being asked.

From the shipped config.json:

ArchitectureGlm5NextForConditionalGeneration, model_type = glm5_next
Layers45 — the first 3 dense, the rest MoE
Hidden dimension4,096
Attention64 query heads, 64 KV heads — full MHA, not grouped
Dense intermediate12,288
Mixture of Experts288 routed experts + 1 shared, 8 routed activated per token, expert intermediate 2,048
Multi-token prediction1 layer
Vocabulary154,880
Native context1,048,576
QuantisationFP8, e4m3, block [128, 128], dynamic activation scaling
LicenceMIT
Context length262,144 tokens
Input modalitiestext only
Streaming✅
Tool calling✅
JSON output✅ json_object
Thinking✅ — on unless you turn it down
Inference enginevLLM
Stabilityexperimental

A system message may sit at any position, and there may be more than one — unlike Qwen3.8-Flash-Next, which allows exactly one and insists it comes first.

Omitting reasoning_effort does not disable thinking. A plain request already returns a populated reasoning. How long the model thinks tracks the prompt rather than the tier — a one-line question may produce a few dozen reasoning tokens, a puzzle several thousand.

Accepted values:

TierNotes
omittedstill thinks
low
medium
highthe longest tier

There is no reasoning_effort value that turns thinking off — none is not in the enum. To keep a response short, cap max_tokens instead, and read the answer from content rather than reasoning.

Where the output lands:

Thinking textchoices[0].message.reasoning
Token countusage.completion_tokens_details.reasoning_tokens

Not available on this endpoint, despite the vision tower in the weights. An image_url part comes back as a 400 whose message describes a server-side path setting rather than the real reason:

Invalid `--allowed-local-media-path`: The path <path> does not exist.

Send text only. For images use DeepSeek-V4-Flash-Vision-Exp, DeepSeek-V4.1-Flash or either Qwen3.8 model.

Context window262,144 tokens, counted as a total budget — prompt plus output, not an output-only allowance
Inputtext only; for images use DeepSeek-V4.1-Flash
JSON outputresponse_format: {"type": "json_object"}
Thinking tiersreasoning_effort (or the equivalent reasoning.effort) — low, medium, high

max_tokens is capped against the total budget, prompt included:

This model's maximum context length is 262144 tokens. However, you requested 128000 output tokens and your prompt contains at least 134145 input tokens, for a total of at least 262145 tokens. Please reduce the length of the input prompt or the number of requested output tokens.

That one is a 400, unlike the 422 above.

Terminal window
curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \
-H "Authorization: Bearer $RADEON_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "GLM-5.3-Flash",
"reasoning_effort": "high",
"messages": [
{ "role": "user", "content": "Why is the sky blue?" }
]
}'

Swap high for xhigh and the same request becomes a 422 on this model — see Thinking.