跳转到内容

Qwen3.8-27B

POST /v1/chat/completions model: Qwen3.8-27B
架构Qwen3_5ForConditionalGeneration,model_type = qwen3_5
层数64
隐藏维度5,120
注意力24 个 Q 头、4 个 KV 头(GQA),头维度 256
中间维度17,408
词表248,320
视觉塔27 层,宽度 1,152,16 头,patch 16,空间合并 2
精度bfloat16——未量化
上下文长度131,072 token
输入模态文本 + 图片
流式✅
工具调用✅
JSON 输出✅ json_object
思考✅——默认开启
推理引擎vLLM
稳定性experimental

不主动调低就一直开着;不传 reasoning_effort 时落到 xhigh。

取值
不传落到 xhigh
low✅
medium✅
xhigh✅ 最长的一档
high minimal max❌ 400

思考文本在 choices[0].message.reasoning。这个模型不返回 usage.completion_tokens_details,思考 token 不单独计数,已包含在 usage.completion_tokens 里。

思考和正文共用同一份 max_tokens 预算。content 返回空时,调大 max_tokens 或传 reasoning_effort: "low"。

system 消息可以不给,但最多一条,且必须放在最前。出现第二条、或排在 user 之后会被拒:

System message must be at the beginning.

角色名写 system 或 developer 都接受。

在 content 里放 image_url 内容块,data: URL 和 https:// URL 都接受。

{
"model": "Qwen3.8-27B",
"messages": [{
"role": "user",
"content": [
{ "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." } },
{ "type": "text", "text": "这张图里写了什么?" }
]
}]
}

图片用量报在 usage.prompt_tokens_details.multimodal_tokens.image。

上下文窗口131,072 token,prompt 与输出合并计算
JSON 输出response_format: {"type": "json_object"}
思考档位low、medium、xhigh

max_tokens 按整个窗口校验:

max_tokens=999999 cannot be greater than max_model_len=max_total_tokens=131072.
Terminal window
curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \
-H "Authorization: Bearer $RADEON_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen3.8-27B",
"reasoning_effort": "low",
"max_tokens": 512,
"messages": [
{ "role": "system", "content": "一句话回答。" },
{ "role": "user", "content": "天为什么是蓝的?" }
]
}'