Qwen3.8-27B
Qwen3.8-27B
Specification
Section titled “Specification”| Architecture | Qwen3_5ForConditionalGeneration, model_type = qwen3_5 |
| Layers | 64 |
| Hidden dimension | 5,120 |
| Attention | 24 query heads, 4 KV heads (GQA), head dim 256 |
| Intermediate dimension | 17,408 |
| Vocabulary | 248,320 |
| Vision tower | 27 layers, width 1,152, 16 heads, patch 16, spatial merge 2 |
| Precision | bfloat16 — unquantised |
At a glance
Section titled “At a glance”| Context length | 131,072 tokens |
| Input modalities | text + images |
| Streaming | ✅ |
| Tool calling | ✅ |
| JSON output | ✅ json_object |
| Thinking | ✅ — on by default |
| Inference engine | vLLM |
| Stability | experimental |
Thinking
Section titled “Thinking”Thinking is on unless you turn it down; omitting reasoning_effort falls through to xhigh.
| Value | |
|---|---|
| omitted | falls through to xhigh |
low | ✅ |
medium | ✅ |
xhigh | ✅ longest tier |
high minimal max | ❌ 400 |
Thinking text arrives in choices[0].message.reasoning. This model does not return
usage.completion_tokens_details, so thinking tokens are not reported separately; they are
included in usage.completion_tokens.
Thinking draws on the same max_tokens budget as the answer. If content comes back empty, raise
max_tokens or send reasoning_effort: "low".
messages
Section titled “messages”A system message is optional, but there can be at most one and it must come first. A second one,
or one after a user turn, is rejected:
System message must be at the beginning.Both system and developer are accepted as the role name.
Image input
Section titled “Image input”Send an image_url content part. Both data: URLs and https:// URLs are accepted.
{ "model": "Qwen3.8-27B", "messages": [{ "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." } }, { "type": "text", "text": "What does this image say?" } ] }]}Image usage is reported at usage.prompt_tokens_details.multimodal_tokens.image.
Limits
Section titled “Limits”| Context window | 131,072 tokens, prompt plus output |
| JSON output | response_format: {"type": "json_object"} |
| Thinking tiers | low, medium, xhigh |
max_tokens is checked against the whole window:
max_tokens=999999 cannot be greater than max_model_len=max_total_tokens=131072.Example
Section titled “Example”curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \ -H "Authorization: Bearer $RADEON_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "Qwen3.8-27B", "reasoning_effort": "low", "max_tokens": 512, "messages": [ { "role": "system", "content": "Answer in one sentence." }, { "role": "user", "content": "Why is the sky blue?" } ] }'