Qwen3.8-27B
POST
/v1/chat/completions
model:
Qwen3.8-27B
| 架构 | Qwen3_5ForConditionalGeneration,model_type = qwen3_5 |
| 层数 | 64 |
| 隐藏维度 | 5,120 |
| 注意力 | 24 个 Q 头、4 个 KV 头(GQA),头维度 256 |
| 中间维度 | 17,408 |
| 词表 | 248,320 |
| 视觉塔 | 27 层,宽度 1,152,16 头,patch 16,空间合并 2 |
| 精度 | bfloat16——未量化 |
| 上下文长度 | 131,072 token |
| 输入模态 | 文本 + 图片 |
| 流式 | ✅ |
| 工具调用 | ✅ |
| JSON 输出 | ✅ json_object |
| 思考 | ✅——默认开启 |
| 推理引擎 | vLLM |
| 稳定性 | experimental |
不主动调低就一直开着;不传 reasoning_effort 时落到 xhigh。
| 取值 | |
|---|---|
| 不传 | 落到 xhigh |
low | ✅ |
medium | ✅ |
xhigh | ✅ 最长的一档 |
high minimal max | ❌ 400 |
思考文本在 choices[0].message.reasoning。这个模型不返回
usage.completion_tokens_details,思考 token 不单独计数,已包含在 usage.completion_tokens 里。
思考和正文共用同一份 max_tokens 预算。content 返回空时,调大 max_tokens
或传 reasoning_effort: "low"。
messages
Section titled “messages”system 消息可以不给,但最多一条,且必须放在最前。出现第二条、或排在 user 之后会被拒:
System message must be at the beginning.角色名写 system 或 developer 都接受。
在 content 里放 image_url 内容块,data: URL 和 https:// URL 都接受。
{ "model": "Qwen3.8-27B", "messages": [{ "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." } }, { "type": "text", "text": "这张图里写了什么?" } ] }]}图片用量报在 usage.prompt_tokens_details.multimodal_tokens.image。
| 上下文窗口 | 131,072 token,prompt 与输出合并计算 |
| JSON 输出 | response_format: {"type": "json_object"} |
| 思考档位 | low、medium、xhigh |
max_tokens 按整个窗口校验:
max_tokens=999999 cannot be greater than max_model_len=max_total_tokens=131072.curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \ -H "Authorization: Bearer $RADEON_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "Qwen3.8-27B", "reasoning_effort": "low", "max_tokens": 512, "messages": [ { "role": "system", "content": "一句话回答。" }, { "role": "user", "content": "天为什么是蓝的?" } ] }'