GLM-5.3-Flash
POST
/v1/chat/completions
model:
GLM-5.3-Flash
关于这个模型
Section titled “关于这个模型”这里跑的权重是 GLM-5.3-Flash,Z.AI 的 GLM-5.3 系列里偏速度的一款—— 稀疏 MoE 模型,回答前先思考,而且不用你开也会思考。
取自随权重发布的 config.json:
| 架构 | Glm5NextForConditionalGeneration,model_type = glm5_next |
| 层数 | 45——前 3 层稠密,其余为 MoE |
| 隐藏维度 | 4,096 |
| 注意力 | 64 个查询头、64 个 KV 头——完整 MHA,不是分组 |
| 稠密中间层 | 12,288 |
| 专家混合 | 288 个路由专家 + 1 个共享,每 token 激活 8 个路由专家,专家中间层 2,048 |
| 多 token 预测 | 1 层 |
| 词表 | 154,880 |
| 原生上下文 | 1,048,576 |
| 量化 | FP8,e4m3,块 [128, 128],动态激活缩放 |
| 许可证 | MIT |
在这个端点上
Section titled “在这个端点上”| 上下文长度 | 262,144 token |
| 输入模态 | 仅文本 |
| 流式 | ✅ |
| 工具调用 | ✅ |
| JSON 输出 | ✅ json_object |
| 思考 | ✅ —— 不主动调低就一直开着 |
| 推理引擎 | vLLM |
| 稳定性 | experimental |
messages
Section titled “messages”system 消息可以放在任意位置,也可以有多条——这一点和
Qwen3.8-Flash-Next 不同,
那个模型只允许一条且必须放在最前面。
不传 reasoning_effort 并不会关掉思考。 普通请求就已经会返回非空的 reasoning。
思考多长取决于提问本身而不是档位——一行问句可能只产生几十个思考 token,
一道谜题则可能用掉几千个。
支持的取值:
| 档位 | 说明 |
|---|---|
| 不传 | 照样思考 |
low | |
medium | |
high | 最高档 |
没有任何一个 reasoning_effort 值能关掉思考——none 不在枚举里。想让回复短一些,
请改用 max_tokens 限制,并从 content 而不是 reasoning 里读答案。
输出位置:
| 思考内容 | choices[0].message.reasoning |
| token 计数 | usage.completion_tokens_details.reasoning_tokens |
这个端点不提供,尽管权重里带着视觉塔。发 image_url 会得到一个 400,
而它的报错描述的是服务端路径配置,并不是真正的原因:
Invalid `--allowed-local-media-path`: The path <path> does not exist.请只发文本。需要图像时请用 DeepSeek-V4-Flash-Vision-Exp、 DeepSeek-V4.1-Flash 或两个 Qwen3.8 模型。
| 上下文窗口 | 262,144 token,是总预算——输入加输出,不是只给输出的额度 |
| 输入 | 仅文本;需要图像请用 DeepSeek-V4.1-Flash |
| JSON 输出 | response_format: {"type": "json_object"} |
| 思考档位 | reasoning_effort(或等价的 reasoning.effort)——low、medium、high |
max_tokens 是对总预算封顶的,输入也算在内:
This model's maximum context length is 262144 tokens. However, you requested 128000 output tokens and your prompt contains at least 134145 input tokens, for a total of at least 262145 tokens. Please reduce the length of the input prompt or the number of requested output tokens.这一条是 400,和上面那个 422 不同。
curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \ -H "Authorization: Bearer $RADEON_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "GLM-5.3-Flash", "reasoning_effort": "high", "messages": [ { "role": "user", "content": "天为什么是蓝的?" } ] }'把 high 换成 xhigh,同样一条请求在这个模型上就是 422——见思考。