Skip to content

DeepSeek-V4-Flash-Vision-Exp

POST /v1/chat/completions model: DeepSeek-V4-Flash-Vision-Exp

A vision tower bolted onto the DeepSeek-V4-Flash weights; the Exp in the name is experimental. It is the only model on this endpoint that has both a one-million-token context and image input.

The language side shares its architecture with the text model, field for field. The table below is read from the config.json of the weights this endpoint actually loads.

ArchitectureDeepseekV4ForCausalLM, model_type = deepseek_v4
Layers43
Hidden size4,096
Attention64 query heads, 1 KV head, head dim 512 (MLA)
Mixture of experts256 routed + 1 shared, 6 activated per token, expert intermediate 2,048
Sparse indexa DSA indexer with index_n_heads / index_topk
Sliding window128
Vocabulary129,280
Native context1,048,576
QuantisationFP8 e4m3, block size [128, 128], dynamic activation scaling, ue8m0 scale format

The vision tower (same config.json, expressed as flat vision_* keys rather than a nested vision_config):

Layers / width / heads32 layers, 1,024 wide, 16 heads
Patch size14
Intermediate size2,816
Downsample ratio3
Per-image token ceiling384
Minimum pixels147,456
Maximum width:height ratio8

vision_max_n_token = 384 is a hard ceiling: a larger image is compressed to 384 tokens, so very high-resolution inputs do not cost linearly more — and do not carry more detail either.

Context length1,048,576 tokens
Input modalitiestext + images
Streaming✅
Tool calling✅
JSON output✅ json_object
Thinking✅ — off by default, ask for it explicitly
Stabilityexperimental
How to sendan {"type":"image_url","image_url":{"url":"data:image/png;base64,…"}} part in content
Meteringusage.prompt_tokens_details.image_tokens
Per-image ceiling384 tokens (above)
Terminal window
curl https://developer.amd.com.cn/radeon/api/v1/chat/completions \
-H "Authorization: Bearer $RADEON_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-Vision-Exp",
"messages": [{
"role": "user",
"content": [
{ "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0KGgo..." } },
{ "type": "text", "text": "What does this image say?" }
]
}]
}'

As permissive as the text model: a system message may sit at any position, there may be more than one, and the role may be spelled either system or developer.

This differs from Qwen3.8-Flash-Next, which accepts only a single leading system. For one client driving both, follow the stricter shape.

Off by default. With reasoning_effort omitted, reasoning_tokens comes back as 0; you have to ask for thinking explicitly.

All six reasoning_effort values are accepted (minimal, low, medium, high, xhigh, max, plus none and omission) — the most permissive model on this endpoint.

Thinking textchoices[0].message.reasoning
Token countusage.completion_tokens_details.reasoning_tokens
Top-level aliasusage.reasoning_tokens — this model does emit it
Context window1,048,576 tokens, counted as prompt plus output
JSON outputresponse_format: {"type": "json_object"}
Turning thinking onreasoning_effort (or the equivalent reasoning.effort)