Skip to content

Qwen3.8-27B

POST /v1/chat/completions model: Qwen3.8-27B
ArchitectureQwen3_5ForConditionalGeneration, model_type = qwen3_5
Layers64
Hidden dimension5,120
Attention24 query heads, 4 KV heads (GQA), head dim 256
Intermediate dimension17,408
Vocabulary248,320
Vision tower27 layers, width 1,152, 16 heads, patch 16, spatial merge 2
Precisionbfloat16 — unquantised
Context length131,072 tokens
Input modalitiestext + images
Streaming✅
Tool calling✅
JSON output✅ json_object
Thinking✅ — on by default
Inference enginevLLM
Stabilityexperimental

Thinking is on unless you turn it down; omitting reasoning_effort falls through to xhigh.

Value
omittedfalls through to xhigh
low✅
medium✅
xhigh✅ longest tier
high minimal max❌ 400

Thinking text arrives in choices[0].message.reasoning. This model does not return usage.completion_tokens_details, so thinking tokens are not reported separately; they are included in usage.completion_tokens.

Thinking draws on the same max_tokens budget as the answer. If content comes back empty, raise max_tokens or send reasoning_effort: "low".

A system message is optional, but there can be at most one and it must come first. A second one, or one after a user turn, is rejected:

System message must be at the beginning.

Both system and developer are accepted as the role name.

Send an image_url content part. Both data: URLs and https:// URLs are accepted.

{
"model": "Qwen3.8-27B",
"messages": [{
"role": "user",
"content": [
{ "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." } },
{ "type": "text", "text": "What does this image say?" }
]
}]
}

Image usage is reported at usage.prompt_tokens_details.multimodal_tokens.image.

Context window131,072 tokens, prompt plus output
JSON outputresponse_format: {"type": "json_object"}
Thinking tierslow, medium, xhigh

max_tokens is checked against the whole window:

max_tokens=999999 cannot be greater than max_model_len=max_total_tokens=131072.
Terminal window
curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \
-H "Authorization: Bearer $RADEON_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen3.8-27B",
"reasoning_effort": "low",
"max_tokens": 512,
"messages": [
{ "role": "system", "content": "Answer in one sentence." },
{ "role": "user", "content": "Why is the sky blue?" }
]
}'