Run locally
DeepSeek-V4-Flash-0731
AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2
- Hardware
- AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2Strix Halo · Radeon 8060S · gfx1151
- Memory
- 128 GiB unified
- Tasks
- Clair Obscure
- Quantisations
- UD-IQ3_XXS
About the model
Mixture-of-Experts model of the DeepSeek-V4 series, released as the official version of DeepSeek-V4-Flash under the MIT license. The repository config.json states 43 layers, a hidden size of 4096, 256 routed experts and one shared expert, of which 6 experts are active per token, 64 attention heads, a vocabulary of 129,280 tokens and max_position_embeddings 1,048,576. The weights are provided in FP8 (format e4m3, block size 128 x 128), and the checkpoint additionally carries a speculative decoding module that the model card calls DSpark and that appears in the configuration as num_nextn_predict_layers 1. The technical report gives 284 billion parameters with 13 billion active parameters for DeepSeek-V4-Flash and a hybrid attention combining Compressed Sparse Attention and Heavily Compressed Attention, while the Hugging Face repository for this version lists 304B params.
- Architektur
- Mixture of Experts, model_type deepseek_v4
- Schichten (num_hidden_layers)
- 43
- Hidden Size
- 4096
- Geroutete Experten
- 256
- Geteilte Experten
- 1
- Aktive Experten je Token
- 6
- Attention-Köpfe
- 64
- Vokabulargröße
- 129 280
Figures from the vendor: Source · Model card · Vendor
The numbers at a glance
Decode, median
7.02 t/s
98 scored responses
Decode, p10 to p90
4.85–9.89 t/s
Antworten ab 200 Tokens werden gewertet; kürzere erzeugen im Protokoll Ausreißer bis über 1 000 000 t/s, weil dort ein einzelnes Token in nahezu null Millisekunden fällt. Median klassisch, Perzentile nach Rangplatz. Beim Prefill gilt diese Schwelle NICHT, dort zählt jede Auswertung.
Peak
10.83 t/s
a single response
Prefill, median
15.2 t/s
Output generated
150,800 Token
scored responses only, from 200 tokens up
Runtime
8.4 h
502 minutes
Self-corrections
12
Responses in total
152
98 of them scored
Input in total
7.81 Mio. Token
the context of all 152 requests added up
of that newly computed
254,590 Token
7.55 million were in the prompt cache
Through the model in total
7.97 Mio. Token
input and output together
At a hosting provider
0.14 USD
DeepSeek-V4-Flash-0731 hosted, cheapest rate
What this run would have cost with DeepSeek-V4-Flash-0731 at a hosting provider
| Provider | Inputper M | Outputper M | Cacheper M | This runwithout cache | This runwith cache |
|---|---|---|---|---|---|
| OpenInference | 0.05 | 0.16 | 0.013 | 0.42 | 0.14 |
| DeepInfra | 0.06 | 0.18 | 0.015 | 0.50 | 0.16 |
| Relace | 0.07 | 0.18 | 0.016 | 0.54 | 0.17 |
| Wafer | 0.10 | 0.25 | 0.050 | 0.82 | 0.44 |
All amounts in USD. The first three columns are tariffs per million tokens, the last two are the total for this run. Gerechnet mit denselben Token wie links: 7 808 346 Eingabe über alle Anfragen, davon 254 590 neu berechnet und 7 553 756 aus dem Prompt-Cache, dazu 158 469 Ausgabe. Die Spalte „ohne Cache“ setzt die ganze Eingabe zum Eingabetarif an, die Spalte „mit Cache“ rechnet den zwischengespeicherten Teil zum Cache-Tarif des Anbieters. Anbieter mit Cache-Tarif sind der Regelfall, deshalb ist die rechte Spalte die realistischere. Es sind die Preise für DeepSeek-V4-Flash-0731 selbst, bei Anbietern die genau dieses Modell hosten, nicht die eines fremden Cloud-Modells. Stand 08.09.2026, Listenpreise ohne Mengenrabatt. Lokal kostet der Lauf statt dessen Strom und die Anschaffung der Karte. Prices at OpenRouter, retrieved 2026-09-08.
What this run would have cost with Sonnet 5 at the same token consumption
| Model | Inputper M | Outputper M | Cache writeper M | Cache readper M | This runwithout cache | This runwith cache |
|---|---|---|---|---|---|---|
| Sonnet 5 | 2.00 | 10.00 | 2.50 | 0.20 | 17.20 | 3.73 |
All amounts in USD. The first four columns are tariffs per million tokens, the last two are the total for this run. Gerechnet mit genau den Token dieses Laufs: 7 808 346 Eingabe, davon 254 590 neu und 7 553 756 aus dem Cache, dazu 158 469 Ausgabe. Ohne Cache zum Eingabetarif, mit Cache zahlt der neue Teil den Schreibtarif und der zwischengespeicherte den Lesetarif. Das ist ein Vergleich unter der Annahme gleichen Tokenverbrauchs. Ein anderes Modell löst dieselbe Aufgabe mit anderer Tokenzahl, meist mit weniger; die Zeile sagt also, was diese Token bei Sonnet 5 gekostet hätten, nicht was Sonnet 5 für diese Aufgabe gekostet hätte. Anthropic price list, retrieved 2026-09-08.
llama.cpp-Serverlog · 98 von 152 Antworten ab 200 Tokens. Der Server lief 506,8 Minuten am Stück. Von der ersten Aufgabe bis zum Ende der letzten sind es 502,3 Minuten, ohne eine einzige Pause über 60 Sekunden dazwischen. Die langen Lücken zwischen Logzeilen sind Rechenblöcke, kein Leerlauf. Die reine Rechenzeit auf der GPU beträgt 500,3 Minuten. · Raw log
Results
Clair Obscur: Expedition 33 — Battle Stage
Opens in its own layerMaus und Tastatur. Angriffe des Gegners werden angekündigt, Parieren und Ausweichen laufen in Echtzeit. Braucht WebGL.
Loads only on click · 0.3 MB transfer · 1.6 MB unpacked
Prefill über Kontexttiefe
A peak figure at short context says little about how a model feels after hours. Hence the curve.
Context depth in thousands of tokens · Momentanrate innerhalb eines einzigen durchgehenden Prefills: Aufgabe 136538 dieses Laufs verarbeitete 36 962 Token am Stück. llama.cpp meldet dabei alle 2 048 Token einen Zwischenstand; die Rate ist der Quotient aus Token und Zeit zwischen zwei aufeinanderfolgenden Ständen. 17 Schritte bleiben übrig, Reststücke unter 1 500 Token sind verworfen, weil dort der Aufruf-Overhead misst statt der Tiefe. Von 2 048 auf 36 446 Token fällt die Rate um den Faktor 2,87. Jeder Punkt ist aus dem verlinkten Serverprotokoll nachrechenbar.
llama.cpp launch parameters
- Context
- 131,072
- KV-Cache
- f16 / f16
- Micro-Batch
- 512
- Speculative
- draft-dspark · n_max 3
- Build
- 580e88d
llama-server -m DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf ^ --alias deepseek-v4-flash --host 127.0.0.1 --port 8100 --device Vulkan0 ^ --gpu-layers all --n-cpu-moe 0 --fit off -fa on --load-mode mmap ^ --ctx-size 131072 --parallel 1 --kv-unified ^ -ctk f16 -ctv f16 -b 2048 -ub 512 ^ --ctx-checkpoints 4 --checkpoint-min-step 4096 ^ -md DSpark-draft.gguf --spec-type draft-dspark --spec-draft-n-max 3 ^ --spec-draft-device Vulkan0 --spec-draft-ngl 99 ^ --jinja --reasoning on --reasoning-format deepseek --reasoning-effort high
Assessment
16 / 40
The grade is a personal, and therefore subjective, assessment.
Game feelDoes it run, and does it play?
4/10
Artefakt liegt vor, zwölf Selbstkorrekturen im Protokoll.
PresentationMenus, HUD, graphics, sound
4/10
Keine Aufnahme vorhanden.
Code qualityTests, structure, self-corrections
4/10
Keine Testdateien.
ScopeHow much of the prompt was fulfilled?
4/10
24 Dateien, 3 901 Zeilen, der kleinste Umfang im Feld.
This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.
How do you see it? The grade above is a personal assessment. Here yours counts, independently of it.
What went wrong
There were 12 self-corrections by the model.
llama.cpp-Serverlog · 98 von 152 Antworten ab 200 Tokens. Der Server lief 506,8 Minuten am Stück. Von der ersten Aufgabe bis zum Ende der letzten sind es 502,3 Minuten, ohne eine einzige Pause über 60 Sekunden dazwischen. Die langen Lücken zwischen Logzeilen sind Rechenblöcke, kein Leerlauf. Die reine Rechenzeit auf der GPU beträgt 500,3 Minuten.
Sources
- Model weightshuggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF
- Model cardhuggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
- Vendor sitedeepseek.com
- Measurement repositorygithub.com/KaiFelixBennett/local-ai-amd-benchmark
- Raw log of this rungithub.com/KaiFelixBennett/local-ai-amd-benchmark/blob/main/evidence/logs/deepseek-v4-flash-clairobscur-halo.log
- Source code · Clair Obscure · UD-IQ3_XXS · Reasoning highgithub.com/KaiFelixBennett/local-ai-amd-benchmark/tree/main/benchmarks/deepseek-v4-flash/clairobscure
Try it yourself
- Download the weights · Source
- Start llama-server with the parameters above
- Register the endpoint in VS Code as a custom model and pick the agent
Prompt, agent files and tools: Agent Test Harness
Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run deepseek-v4-flash-clairobscure-evox2. Measurement data CC-BY-4.0.
Measurement data CC-BY-4.0, code MIT.