Skip to content

Run locally

DeepSeek-V4-Flash-0731

AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2

Hardware
AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2Strix Halo · Radeon 8060S · gfx1151
Memory
128 GiB unified
Tasks
Clair Obscure
Quantisations
UD-IQ3_XXS

About the model

Mixture-of-Experts model of the DeepSeek-V4 series, released as the official version of DeepSeek-V4-Flash under the MIT license. The repository config.json states 43 layers, a hidden size of 4096, 256 routed experts and one shared expert, of which 6 experts are active per token, 64 attention heads, a vocabulary of 129,280 tokens and max_position_embeddings 1,048,576. The weights are provided in FP8 (format e4m3, block size 128 x 128), and the checkpoint additionally carries a speculative decoding module that the model card calls DSpark and that appears in the configuration as num_nextn_predict_layers 1. The technical report gives 284 billion parameters with 13 billion active parameters for DeepSeek-V4-Flash and a hybrid attention combining Compressed Sparse Attention and Heavily Compressed Attention, while the Hugging Face repository for this version lists 304B params.

Architecture
Mixture of Experts, model_type deepseek_v4
Layers (num_hidden_layers)
43
Hidden size
4096
Routed experts
256
Shared experts
1
Active experts per token
6
Attention heads
64
Vocabulary size
129,280

Figures from the vendor: Source · Model card · Vendor

The numbers at a glance

Decode, median

7.02 t/s

98 scored responses

Decode, p10 to p90

4.85–9.89 t/s

Responses from 200 tokens up are scored; shorter ones produce outliers of over 1,000,000 t/s in the log, because a single token there takes almost zero milliseconds. Classic median, percentiles by rank. This threshold does NOT apply to prefill, where every evaluation counts.

Peak

10.83 t/s

a single response

Prefill, median

15.2 t/s

Output generated

150,800 Token

scored responses only, from 200 tokens up

Runtime

8.4 h

502 minutes

Self-corrections

12

Responses in total

152

98 of them scored

Input in total

7.81 Mio. Token

the context of all 152 requests added up

of that newly computed

254,590 Token

7.55 million were in the prompt cache

Through the model in total

7.97 Mio. Token

input and output together

At a hosting provider

0.14 USD

DeepSeek-V4-Flash-0731 hosted, cheapest rate

What this run would have cost with DeepSeek-V4-Flash-0731 at a hosting provider
Provider Inputper M Outputper M Cacheper M This runwithout cache This runwith cache
OpenInference 0.050.16 0.013 0.42 0.14
DeepInfra 0.060.18 0.015 0.50 0.16
Relace 0.070.18 0.016 0.54 0.17
Wafer 0.100.25 0.050 0.82 0.44

All amounts in USD. The first three columns are tariffs per million tokens, the last two are the total for this run. Calculated with the same tokens as on the left: 7,808,346 input across all requests, of which 254,590 were newly computed and 7,553,756 came from the prompt cache, plus 158,469 output. The “without cache” column charges the entire input at the input rate; the “with cache” column charges the cached part at the provider's cache rate. Providers with a cache rate are the norm, so the right-hand column is the more realistic one. These are the prices for DeepSeek-V4-Flash-0731 itself, at providers that host exactly this model, not those of some other cloud model. As of 8 September 2026, list prices without volume discounts. Locally, the run costs electricity and the purchase of the card instead. Prices at OpenRouter, retrieved 2026-09-08.

What this run would have cost with Sonnet 5 at the same token consumption
Model Inputper M Outputper M Cache writeper M Cache readper M This runwithout cache This runwith cache
Sonnet 5 2.0010.00 2.500.20 17.20 3.73

All amounts in USD. The first four columns are tariffs per million tokens, the last two are the total for this run. Calculated with exactly the tokens of this run: 7,808,346 input, of which 254,590 new and 7,553,756 from the cache, plus 158,469 output. Without cache at the input rate; with cache, the new part pays the write rate and the cached part the read rate. This is a comparison that assumes the same token consumption. A different model solves the same task with a different number of tokens, usually fewer; so the row says what these tokens would have cost with Sonnet 5, not what Sonnet 5 would have cost for this task. Anthropic price list, retrieved 2026-09-08.

llama.cpp server log · 98 of 152 responses from 200 tokens up. The server ran for 506.8 minutes in one go. From the first task to the end of the last it is 502.3 minutes, without a single pause over 60 seconds in between. The long gaps between log lines are compute blocks, not idle time. The pure compute time on the GPU is 500.3 minutes. · Raw log

Results

Clair Obscur: Expedition 33 — Battle Stage

Opens in its own layer

Mouse and keyboard. Enemy attacks are announced, parrying and dodging happen in real time. Needs WebGL.

Loads only on click · 0.3 MB transfer · 1.6 MB unpacked

Prefill over context depth

A peak figure at short context says little about how a model feels after hours. Hence the curve.

184470 3.1 K · 65.3 t/s 9.2 K · 43.8 t/s 13.3 K · 36.2 t/s 19.5 K · 29.8 t/s 25.6 K · 32.0 t/s 29.7 K · 25.7 t/s 35.6 K · 22.7 t/s 3.1 K35.6 K t/s

Context depth in thousands of tokens · Instantaneous rate within a single continuous prefill: task 136538 of this run processed 36,962 tokens in one go. llama.cpp reports an intermediate status every 2,048 tokens; the rate is the quotient of tokens and time between two consecutive reports. 17 steps remain; leftover pieces under 1,500 tokens are discarded, because there the call overhead is measured rather than depth. From 2,048 to 36,446 tokens the rate drops by a factor of 2.87. Every point can be recalculated from the linked server log.

llama.cpp launch parameters

Context
131,072
KV-Cache
f16 / f16
Micro-Batch
512
Speculative
draft-dspark · n_max 3
Build
580e88d
llama-server -m DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf ^
  --alias deepseek-v4-flash --host 127.0.0.1 --port 8100 --device Vulkan0 ^
  --gpu-layers all --n-cpu-moe 0 --fit off -fa on --load-mode mmap ^
  --ctx-size 131072 --parallel 1 --kv-unified ^
  -ctk f16 -ctv f16 -b 2048 -ub 512 ^
  --ctx-checkpoints 4 --checkpoint-min-step 4096 ^
  -md DSpark-draft.gguf --spec-type draft-dspark --spec-draft-n-max 3 ^
  --spec-draft-device Vulkan0 --spec-draft-ngl 99 ^
  --jinja --reasoning on --reasoning-format deepseek --reasoning-effort high

Assessment

16 / 40

The grade is a personal, and therefore subjective, assessment.

Game feelDoes it run, and does it play?

Artefact available, twelve self-corrections in the log.

PresentationMenus, HUD, graphics, sound

No recording available.

Code qualityTests, structure, self-corrections

No test files.

ScopeHow much of the prompt was fulfilled?

24 files, 3,901 lines, the smallest scope in the field.

This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.

What went wrong

There were 12 self-corrections by the model.

llama.cpp server log · 98 of 152 responses from 200 tokens up. The server ran for 506.8 minutes in one go. From the first task to the end of the last it is 502.3 minutes, without a single pause over 60 seconds in between. The long gaps between log lines are compute blocks, not idle time. The pure compute time on the GPU is 500.3 minutes.

Try it yourself

  1. Download the weights · Source
  2. Start llama-server with the parameters above
  3. Register the endpoint in VS Code as a custom model and pick the agent

Prompt, agent files and tools: Agent Test Harness

Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run deepseek-v4-flash-clairobscure-evox2. Measurement data CC-BY-4.0.

Measurement data CC-BY-4.0, code MIT.