Run locally
Qwen3.5-122B-A10B
AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2
- Hardware
- AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2Strix Halo · Radeon 8060S · gfx1151
- Memory
- 128 GiB unified
- Tasks
- Labortest, Image recognition
- Quantisations
- UD-Q4_K_XL
About the model
Sparse Mixture-of-Experts with 122 billion total parameters and 10 billion active per token. The model card states 48 layers in the pattern 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE)), that is three Gated DeltaNet layers per one Gated Attention layer. Each MoE layer holds 256 experts, of which 8 routed plus 1 shared expert are active. The card gives the type as "Causal Language Model with Vision Encoder"; native context length is 262,144 tokens and, per the card, extensible to 1,010,000 tokens.
- Type
- Causal Language Model with Vision Encoder (Sparse MoE)
- Total parameters
- 122 B
- Active parameters per token
- 10 B
- Layers
- 48
- Experts per MoE layer
- 256
- Active experts
- 8 routed + 1 shared
- Context length
- 262,144 tokens native, extensible up to 1,010,000 tokens
- License
- apache-2.0
Figures from the vendor: Source · Model card · Vendor
The numbers at a glance
Synthetic lab test with pp512 and tg128, no agent run. The speed therefore comes from a measurement series, not from scored responses of a real run. There is no token balance and no curve over context depth; both require many consecutive requests. The artefact next to it comes from a separate run of the same model.
Decode with MTP
31.80 t/s
draft depth 2, a single prompt, 1.55 times the value without speculation
Decode on the tasks
26.5 up to 28.0 t/s
with MTP, draft depth 2, across the five programming tasks
Decode without speculation
20.52 t/s
llama-bench tg128
Draft acceptance
0.866
at draft depth 2, on average 2.73 tokens per step
Prefill at 512 tokens
245.71 t/s
llama-bench pp512
Prefill at 4,096 tokens
232.68 t/s
only 5 percent below the value at 512 tokens
Programming tasks
5 of 5
run against hidden tests, 1,232 seconds in total
Weights
73.23 GiB
UD-Q4_K_XL, 124.64 billion parameters, about 10 billion active
Synthetic: llama-bench pp512, pp4096 and tg128, plus llama-server for the speculation. Measurement series of 23 July 2026. · Raw log
Results
The game this run delivered does not start: three syntax errors abort loading, leaving a black screen. That is why there is neither a recording nor a version to try out here. The source code is in the evidence repo exactly as the model handed it in.
Measurement series
Speed over draft depth
MTP predicts several tokens at once, and the model checks them in one step. How many it tries per step is set by --spec-draft-n-max. More is not automatically faster: at 6 so much is discarded that the gain shrinks.
llama.cpp launch parameters
- Context
- 262,144
- KV-Cache
- q8_0 / q8_0
- Micro-Batch
- 512
- Speculative
- draft-mtp · n_max 2 · Annahme 0,866
- Build
- 04b2b72
llama-server -m Qwen3.5-122B-A10B-UD-Q4_K_XL-00001-of-00003.gguf ^ --mmproj mmproj-F16.gguf --host 127.0.0.1 --port 8098 ^ -c 262144 -ngl 999 -fa on ^ -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 ^ --spec-type draft-mtp --spec-draft-n-max 2 ^ --temp 0.6 --top-p 0.95 --presence-penalty 1.5 ^ --jinja --chat-template-file qwen-fixed-v21.3.jinja ^ --reasoning on --reasoning-format deepseek
Assessment
4 / 40
The grade is a personal, and therefore subjective, assessment.
Game feelDoes it run, and does it play?
0/10
The game does not start. Three syntax errors abort loading, leaving a black screen.
PresentationMenus, HUD, graphics, sound
0/10
None of it can be seen: without a working build there are no menus, no graphics, no sound.
Code qualityTests, structure, self-corrections
2/10
No tests, and three syntax errors that a single start would have shown immediately.
ScopeHow much of the prompt was fulfilled?
2/10
A lot is laid out: combat system, reactions, target selection, turn order, sound and post-processing in 6,537 lines. None of it is delivered in the result, because the game does not start.
This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.
How do you see it? The grade above is a personal assessment. Here yours counts, independently of it.
What went wrong
- The delivered game does not start. The model rewrote three loops from forEach to for and each time left the old }); in place: in damage-numbers.js line 142, hud.js line 263 and prompts.js line 274. Each of these spots aborts loading of the whole game, leaving a black screen.
- With Unsloth's preset for precise coding, presence_penalty 0.0, the model never finished one task: 32,768 tokens over 1,087 seconds, without an answer. Laguna solves the same task in 201 tokens. With presence_penalty 1.5 at temperature 0.6 it reaches the goal with 7,680 tokens in 297 seconds. The first pass of the programming tasks is therefore invalid.
- Unsloth's example value for the draft depth, 6, is 13 percent below the measured optimum.
- Same result as Laguna on the programming tasks, but 67 percent slower: the model thinks for 3,000 to 6,600 tokens on every task, even the clearly specified ones, and so produces 2.3 times as many tokens.
- Not checked is whether MTP holds up under real agent load. With Laguna, draft acceptance collapsed to zero exactly there. The 31.80 t/s come from a single prompt.
- Not measured yet: speed over context depth and tool calling.
- On the second image task the model could not read a single price tag. It said so itself and invented nothing.
Synthetic: llama-bench pp512, pp4096 and tg128, plus llama-server for the speculation. Measurement series of 23 July 2026.
Image recognition
A separate test run with the same model on the same machine, not part of the run in which the game was built.
Values read
30 of 50
dashboard, 50 checkable entries
Price tags
0 of 4
read fully correctly
Invented tags
0
none claimed as certain
Stumbling points spotted
1 of 4
places that break the pattern
Provisional: the answer key for the price tags is not fully confirmed yet.
Task 1 · reading values
What the model read from the tooltip table
| Claude Opus 4.5 | Qwen3.5 27B | |
|---|---|---|
| Intelligence Index | 40.8 | 37.1 |
| GPQA diamond | 87 | 84 |
| Humanity's Last Exam | 29 | 22 |
| Terminal-bench 2.0 | 50 | 61 |
And from the KPI tiles, in German as on the test image
- QWEN3.5 27B 37.1 · AI Intelligence Index — Open-Weight, kein API-Key, läuft auf der eigenen GPU
- VS. OPUS 4.5 (49.7) 91 · des Opus-4.5-Niveaus im Gesamtindex. Bei Coding-Benchmarks: gleichauf.
- PROPRIETÄRE SPITZE · HEUTE 59.9 · Das Rechenzentrum klettert weiter — doch der Abstand schrumpft jedes Quartal.
Task 2 · complex text recognition
What is tested is the reading of tiny, tilted, partly hidden text in a photo with more than forty objects, plus two currency symbols that are hard to tell apart once rotated. The price tags are what the test is run on.
What the model read
No tags named.
The model on the hardest part, quoted in the German it answered in: „Die Preisschilder im unteren linken Bereich (Wertungsbereich: x<700, y>950) sind extrem unscharf, gedreht und teilweise verdeckt. Keine einzige Zahl war sicher lesbar.“
Where the model called itself unsure (3 places)
- Alle Preisschilder in Bild B - keine Zahlen konnten sicher gelesen werden
- Genaue Positionsbestimmung welches Schild im Wertungsbereich liegt bei dieser Perspektive
- Unterscheidung von £ und € bei gedrehten Schildern
The model wrote this list itself, not the scorer, and in German; it is quoted as written. It counts: an admitted “not readable” costs one point, an invented entry costs more.
The complete answer of the model as JSON · The scoring is not published beside it: it names the correct value for every mistake, which would put the answer key in the open.
Sources
- Model weightshuggingface.co/unsloth/Qwen3.5-122B-A10B-GGUF
- Model cardhuggingface.co/Qwen/Qwen3.5-122B-A10B
- Vendor siteqwen.ai
- Measurement repositorygithub.com/KaiFelixBennett/local-ai-amd-benchmark
- Raw log of this rungithub.com/KaiFelixBennett/local-ai-amd-benchmark/blob/main/evidence/reports/qwen35-122b-a10b-strix-halo-vulkan-benchmark.md
- Source code · Labortest · UD-Q4_K_XL · Reasoning on, no limitgithub.com/KaiFelixBennett/local-ai-amd-benchmark/tree/main/benchmarks/qwen35-122b-a10b/clairobscure
Try it yourself
- Download the weights · Source
- Start llama-server with the parameters above
- Register the endpoint in VS Code as a custom model and pick the agent
Prompt, agent files and tools: Agent Test Harness
Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run qwen35-122b-a10b-evox2. Measurement data CC-BY-4.0.
Measurement data CC-BY-4.0, code MIT.