Run locally
Qwen3.8-Flash-Next
AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2 · 2 runs
- Hardware
- AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2Strix Halo · Radeon 8060S · gfx1151
- Memory
- 128 GiB unified
- Tasks
- Moorhuhn, Clair Obscure, Image recognition
- Quantisations
- UD-Q4_K_XL
Personal opinion
Kai Bennett
Until now I had not found a model that works really well on the AMD Strix Halo, because either they were dense models that ran too slowly with the low bandwidth, or they were so small that they also ran well on hardware with much less memory, or they had, for example, no image recognition capabilities.
This model changes that. In the benchmarks it can keep up with much larger models and is even ranked above GPT 5.6 Luna (max) at artificialanalysis.ai, it runs consistently at around 20 tokens even with larger context, has vision capabilities, uses the full memory of the device and delivers good results.
In my view, the model is practically made for the AMD Strix Halo or DGX Spark.
About the model
Hybrid MoE model with 125 billion parameters in the language model, 6 billion of them active per token, plus 51 billion parameters of n-gram embedding and 4 billion MTP. The model card states 48 layers in the pattern 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE)) and 512 experts, of which 10 routed and 1 shared are active per token. Native context is 262,144 tokens and can be extended to 1,000,000 tokens. The type is given as "Causal Language Model with Vision Encoder", and config.json contains a vision encoder with 27 layers and hidden size 1152.
- Total parameters, language model
- 125B
- Active parameters per token
- 6B
- N-gram embedding
- 51B
- MTP
- 4B, 1 layer
- Layers
- 48
- Layer pattern
- 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
- Experts
- 512
- Active experts
- 10 routed + 1 shared
- Context length
- 262,144 native, extensible to 1,000,000 tokens
- License
- other, license_name qwen-community-1.0
Figures from the vendor: Source · Model card · Vendor
Run: Moorhuhn · UD-Q4_K_XL · reasoning xhigh · 03 September 2026
The numbers at a glance
Decode, median
21.76 t/s
346 scored responses
Decode, p10 to p90
17.15–26.28 t/s
Responses from 200 tokens up are scored; shorter ones produce outliers of over 1,000,000 t/s in the log, because a single token there takes almost zero milliseconds. Classic median, percentiles by rank. This threshold does NOT apply to prefill, where every evaluation counts.
Peak
33.89 t/s
a single response
Prefill, median
77.6 t/s
Output generated
770,428 Token
scored responses only, from 200 tokens up
Runtime
13.6 h
815 minutes
Self-corrections
not measured
Responses in total
459
346 of them scored
Input in total
24.15 Mio. Token
the context of all 459 requests added up
of that newly computed
1,051,292 Token
23.10 million were in the prompt cache
Through the model in total
24.94 Mio. Token
input and output together
At a hosting provider
0.95 USD
Qwen3.8 Flash hosted, cheapest rate
What this run would have cost with Qwen3.8 Flash at a hosting provider
| Provider | Inputper M | Outputper M | Cacheper M | This runwithout cache | This runwith cache |
|---|---|---|---|---|---|
| Alibaba Cloud Int. | 0.15 | 0.47 | 0.016 | 3.99 | 0.95 |
All amounts in USD. The first three columns are tariffs per million tokens, the last two are the total for this run. Calculated with the same tokens as on the left: 24,148,766 input across all requests, of which 1,051,292 were newly computed and 23,097,474 came from the prompt cache, plus 786,506 output. The “without cache” column charges the entire input at the input rate. In the “with cache” column, the new part pays the write rate of 0.20 USD and the cached part the read rate of 0.016 USD. The only provider is Alibaba Cloud International. There the model is called “Qwen3.8 Flash”, without the suffix Next and without a parameter count. What ran here is the Qwen3.8-Flash-Next build from Hugging Face with 125 billion parameters, 6 billion of them active. The same vendor, the same context length of 1,000,000 tokens and the same release date point to the same origin; the provider page does not prove it. As of 8 September 2026, list prices without volume discounts. Locally, the run costs electricity and the purchase of the computer instead. Prices at OpenRouter, retrieved 2026-09-08.
What this run would have cost with Sonnet 5 at the same token consumption
| Model | Inputper M | Outputper M | Cache writeper M | Cache readper M | This runwithout cache | This runwith cache |
|---|---|---|---|---|---|---|
| Sonnet 5 | 2.00 | 10.00 | 2.50 | 0.20 | 56.16 | 15.11 |
All amounts in USD. The first four columns are tariffs per million tokens, the last two are the total for this run. Calculated with exactly the tokens of this run: 24,148,766 input, of which 1,051,292 new and 23,097,474 from the cache, plus 786,506 output. Without cache at the input rate; with cache, the new part pays the write rate and the cached part the read rate. This is a comparison that assumes the same token consumption. A different model solves the same task with a different number of tokens, usually fewer; so the row says what these tokens would have cost with Sonnet 5, not what Sonnet 5 would have cost for this task. Anthropic price list, retrieved 2026-09-08.
llama.cpp server log · 346 of 459 responses from 200 tokens up · run of 3/4 September 2026, 13 h 35 min. The server ran for 844.3 minutes in one go. From the first task to the end of the last one it is 837.0 minutes; this includes one idle pause of 21.9 minutes. The working window is therefore 815 minutes. Pure compute time on the GPU is 798.8 minutes. · Raw log
Results
Moorland Mayhem — Featherstorm
Opens in its own layerAim and shoot with mouse or finger. Reloading happens automatically in five of seven modes. Needs WebGL. In portrait mode the targets are smaller than a fingertip, so play in landscape in full screen.
Loads only on click · 0.3 MB transfer · 1.3 MB unpacked
Prefill over context depth
A peak figure at short context says little about how a model feels after hours. Hence the curve.
Context depth in thousands of tokens · Instantaneous rate within a single continuous prefill: task 185882 of this run processed 69,985 tokens in one go. llama.cpp reports an intermediate status every 2,048 tokens; the rate is the quotient of tokens and time between two consecutive reports. 32 steps remain, values under 1,500 tokens are excluded. From 2,048 to 68,827 tokens the rate drops by a factor of 1.58. Every point can be traced in the linked server log.
llama.cpp launch parameters
- Context
- 131,072
- KV-Cache
- q8_0 / q8_0
- Micro-Batch
- 512
- Speculative
- MTP, Shared-Q8_0, n-max 2
- Build
- —
llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf ^ --mmproj mmproj-F16.gguf --alias qwen3.8-flash-next ^ --host 127.0.0.1 --port 8099 --device Vulkan0 ^ --gpu-layers all --n-cpu-moe 0 --fit off -fa on ^ --load-mode mmap --lazy-mode on ^ --ctx-size 262144 --parallel 1 --kv-unified ^ -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 ^ --ctx-checkpoints 4 --checkpoint-min-step 4096 ^ --jinja --reasoning on --reasoning-format deepseek --reasoning-effort xhigh
Assessment
35 / 40
The grade is a personal, and therefore subjective, assessment.
Game feelDoes it run, and does it play?
9/10
A very good game feel and feedback, playing is fun and varied.
PresentationMenus, HUD, graphics, sound
9/10
Complete implementations, menus are nicely and extensively designed. Best result, better than Sonnet 5.
Code qualityTests, structure, self-corrections
8/10
12 test files with 1,494 test lines, no repair script. Sensible folder structure, it also created a subfolder.
ScopeHow much of the prompt was fulfilled?
9/10
55 files, 10,531 lines. Almost everything from the prompt was actually implemented in the game.
How do you see it? The grade above is a personal assessment. Here yours counts, independently of it.
What went wrong
- The loading screen says “Featherstorm wordt geladen”, Dutch instead of German. Exactly one spot in the whole game; the rest is German throughout.
- The game renders to a 1280×720 canvas. On a phone 390 pixels wide it is scaled down to 390×219, playable, but the targets are small.
- The run needed three sessions (20, 376 and 419 minutes). Between the sessions the context was lost and had to be rebuilt.
llama.cpp server log · 346 of 459 responses from 200 tokens up · run of 3/4 September 2026, 13 h 35 min. The server ran for 844.3 minutes in one go. From the first task to the end of the last one it is 837.0 minutes; this includes one idle pause of 21.9 minutes. The working window is therefore 815 minutes. Pure compute time on the GPU is 798.8 minutes.
Run: Clair Obscure · UD-Q4_K_XL · reasoning xhigh · 01 September 2026
The numbers at a glance
Decode, median
10.93 t/s
92 scored responses
Decode, p10 to p90
9.70–13.85 t/s
Responses from 200 tokens up are scored; shorter ones produce outliers of over 1,000,000 t/s in the log, because a single token there takes almost zero milliseconds. Classic median, percentiles by rank. This threshold does NOT apply to prefill, where every evaluation counts.
Peak
17.01 t/s
a single response
Prefill, median
109.7 t/s
Output generated
236,460 Token
scored responses only, from 200 tokens up
Runtime
6.2 h
370 minutes
Self-corrections
not measured
Responses in total
106
92 of them scored
Input in total
8.15 Mio. Token
the context of all 106 requests added up
of that newly computed
251,883 Token
7.90 million were in the prompt cache
Through the model in total
8.39 Mio. Token
input and output together
At a hosting provider
0.29 USD
Qwen3.8 Flash hosted, cheapest rate
What this run would have cost with Qwen3.8 Flash at a hosting provider
| Provider | Inputper M | Outputper M | Cacheper M | This runwithout cache | This runwith cache |
|---|---|---|---|---|---|
| Alibaba Cloud Int. | 0.15 | 0.47 | 0.016 | 1.34 | 0.29 |
All amounts in USD. The first three columns are tariffs per million tokens, the last two are the total for this run. Calculated with the same tokens as on the left: 8,153,803 input across all requests, of which 251,883 were newly computed and 7,901,920 came from the prompt cache, plus 238,371 output. The “without cache” column charges the entire input at the input rate. In the “with cache” column, the new part pays the write rate of 0.20 USD and the cached part the read rate of 0.016 USD. The only provider is Alibaba Cloud International. There the model is called “Qwen3.8 Flash”, without the suffix Next and without a parameter count. What ran here is the Qwen3.8-Flash-Next build from Hugging Face with 125 billion parameters, 6 billion of them active. The same vendor, the same context length of 1,000,000 tokens and the same release date point to the same origin; the provider page does not prove it. As of 8 September 2026, list prices without volume discounts. Locally, the run costs electricity and the purchase of the computer instead. Prices at OpenRouter, retrieved 2026-09-08.
What this run would have cost with Sonnet 5 at the same token consumption
| Model | Inputper M | Outputper M | Cache writeper M | Cache readper M | This runwithout cache | This runwith cache |
|---|---|---|---|---|---|---|
| Sonnet 5 | 2.00 | 10.00 | 2.50 | 0.20 | 18.69 | 4.59 |
All amounts in USD. The first four columns are tariffs per million tokens, the last two are the total for this run. Calculated with exactly the tokens of this run: 8,153,803 input, of which 251,883 new and 7,901,920 from the cache, plus 238,371 output. Without cache at the input rate; with cache, the new part pays the write rate and the cached part the read rate. This is a comparison that assumes the same token consumption. A different model solves the same task with a different number of tokens, usually fewer; so the row says what these tokens would have cost with Sonnet 5, not what Sonnet 5 would have cost for this task. Anthropic price list, retrieved 2026-09-08.
llama.cpp server log · 92 of 106 responses from 200 tokens up · run from August 2026. The server ran for 374.4 minutes in one go. From the first task to the end of the last one it is 369.7 minutes, without a single pause of more than 60 seconds in between. Pure compute time on the GPU is 368.1 minutes. · Raw log
Results
Clair Obscur Echo — Studio of the Gilded Hour
Opens in its own layerMouse and keyboard. Enemy attacks are announced, parrying and dodging happen in real time. Needs WebGL.
Loads only on click · 0.3 MB transfer · 1.7 MB unpacked
Prefill over context depth
A peak figure at short context says little about how a model feels after hours. Hence the curve.
Context depth in thousands of tokens · Instantaneous rate within a single continuous prefill: task 176155 of this run processed 35,009 tokens in one go. llama.cpp reports an intermediate status every 2,048 tokens; the rate is the quotient of tokens and time between two consecutive reports. 15 steps remain; leftover pieces under 1,500 tokens are discarded, because there the call overhead is measured rather than depth. From 2,048 to 33,642 tokens the rate drops by a factor of 1.19. Every point can be recalculated from the linked server log.
llama.cpp launch parameters
- Context
- 262,144
- KV-Cache
- q8_0 / q8_0
- Micro-Batch
- 512
- Speculative
- MTP, older version (not the one from Unsloth)
- Build
- 580e88d
llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf ^ --mmproj mmproj-F16.gguf --alias qwen3.8-flash-next ^ --host 127.0.0.1 --port 8099 --device Vulkan0 ^ --gpu-layers all --n-cpu-moe 0 --fit off -fa on ^ --load-mode mmap --lazy-mode on ^ --ctx-size 262144 --parallel 1 --kv-unified ^ -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 ^ --ctx-checkpoints 4 --checkpoint-min-step 4096 ^ --jinja --reasoning on --reasoning-format deepseek --reasoning-effort xhigh
Assessment
22 / 40
The grade is a personal, and therefore subjective, assessment.
Game feelDoes it run, and does it play?
6/10
Combat arena runs, parrying and dodging are in place.
PresentationMenus, HUD, graphics, sound
6/10
Recording available.
Code qualityTests, structure, self-corrections
4/10
No test files; no repair scripts either.
ScopeHow much of the prompt was fulfilled?
6/10
29 files, 6,157 lines.
This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.
How do you see it? The grade above is a personal assessment. Here yours counts, independently of it.
What went wrong
No self-corrections were logged.
llama.cpp server log · 92 of 106 responses from 200 tokens up · run from August 2026. The server ran for 374.4 minutes in one go. From the first task to the end of the last one it is 369.7 minutes, without a single pause of more than 60 seconds in between. Pure compute time on the GPU is 368.1 minutes.
Sources
- Model weightshuggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
- Model cardhuggingface.co/Qwen/Qwen3.8-Flash-Next
- Vendor siteqwen.ai
- Measurement repositorygithub.com/KaiFelixBennett/local-ai-amd-benchmark
- Raw log of this rungithub.com/KaiFelixBennett/local-ai-amd-benchmark/blob/main/evidence/logs/qwen38-flashnext-moorhuhn-evox2.log
- Source code · Moorhuhn · UD-Q4_K_XL · Reasoning xhighgithub.com/KaiFelixBennett/local-ai-amd-benchmark/tree/main/benchmarks/qwen38-flashnext/moorhuhn-q4xl
- Source code · Clair Obscure · UD-Q4_K_XL · Reasoning xhighgithub.com/KaiFelixBennett/local-ai-amd-benchmark/tree/main/benchmarks/qwen38-flashnext/clairobscure-q4xl
Try it yourself
- Download the weights · Source
- Start llama-server with the parameters above
- Register the endpoint in VS Code as a custom model and pick the agent
Prompt, agent files and tools: Agent Test Harness
Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run qwen38-flashnext-moorhuhn-evox2. Measurement data CC-BY-4.0.
Measurement data CC-BY-4.0, code MIT.