Skip to content

Run locally

Qwen3.5-122B-A10B

AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2

Hardware
AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2Strix Halo · Radeon 8060S · gfx1151
Memory
128 GiB unified
Tasks
Labortest, Image recognition
Quantisations
UD-Q4_K_XL

About the model

Sparse Mixture-of-Experts with 122 billion total parameters and 10 billion active per token. The model card states 48 layers in the pattern 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE)), that is three Gated DeltaNet layers per one Gated Attention layer. Each MoE layer holds 256 experts, of which 8 routed plus 1 shared expert are active. The card gives the type as "Causal Language Model with Vision Encoder"; native context length is 262,144 tokens and, per the card, extensible to 1,010,000 tokens.

Type
Causal Language Model with Vision Encoder (Sparse MoE)
Total parameters
122 B
Active parameters per token
10 B
Layers
48
Experts per MoE layer
256
Active experts
8 routed + 1 shared
Context length
262,144 tokens native, extensible up to 1,010,000 tokens
License
apache-2.0

Figures from the vendor: Source · Model card · Vendor

The numbers at a glance

Synthetic lab test with pp512 and tg128, no agent run. The speed therefore comes from a measurement series, not from scored responses of a real run. There is no token balance and no curve over context depth; both require many consecutive requests. The artefact next to it comes from a separate run of the same model.

Decode with MTP

31.80 t/s

draft depth 2, a single prompt, 1.55 times the value without speculation

Decode on the tasks

26.5 up to 28.0 t/s

with MTP, draft depth 2, across the five programming tasks

Decode without speculation

20.52 t/s

llama-bench tg128

Draft acceptance

0.866

at draft depth 2, on average 2.73 tokens per step

Prefill at 512 tokens

245.71 t/s

llama-bench pp512

Prefill at 4,096 tokens

232.68 t/s

only 5 percent below the value at 512 tokens

Programming tasks

5 of 5

run against hidden tests, 1,232 seconds in total

Weights

73.23 GiB

UD-Q4_K_XL, 124.64 billion parameters, about 10 billion active

Synthetic: llama-bench pp512, pp4096 and tg128, plus llama-server for the speculation. Measurement series of 23 July 2026. · Raw log

Results

The game this run delivered does not start: three syntax errors abort loading, leaving a black screen. That is why there is neither a recording nor a version to try out here. The source code is in the evidence repo exactly as the model handed it in.

Measurement series

Speed over draft depth

MTP predicts several tokens at once, and the model checks them in one step. How many it tries per step is set by --spec-draft-n-max. More is not automatically faster: at 6 so much is discarded that the gain shrinks.

192633 0 · 20.55 t/s · without speculation 2 · 31.80 t/s · acceptance 0.866, length 2.73 3 · 31.05 t/s · acceptance 0.726, length 3.18 4 · 31.89 t/s · acceptance 0.725, length 3.90 6 · 27.68 t/s · acceptance 0.530, length 4.18 in operationUnsloth's example value 02346 t/s
Draft depth, --spec-draft-n-max · One run per depth, freshly loaded server, the same prompt, KV q8_0, thinking on. 2 and 4 are 0.3 percent apart and count as equally fast. 2 was chosen because its acceptance of 0.866 leaves more headroom. Report, section 5.3.

llama.cpp launch parameters

Context
262,144
KV-Cache
q8_0 / q8_0
Micro-Batch
512
Speculative
draft-mtp · n_max 2 · Annahme 0,866
Build
04b2b72
llama-server -m Qwen3.5-122B-A10B-UD-Q4_K_XL-00001-of-00003.gguf ^
  --mmproj mmproj-F16.gguf --host 127.0.0.1 --port 8098 ^
  -c 262144 -ngl 999 -fa on ^
  -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 ^
  --spec-type draft-mtp --spec-draft-n-max 2 ^
  --temp 0.6 --top-p 0.95 --presence-penalty 1.5 ^
  --jinja --chat-template-file qwen-fixed-v21.3.jinja ^
  --reasoning on --reasoning-format deepseek

Assessment

4 / 40

The grade is a personal, and therefore subjective, assessment.

Game feelDoes it run, and does it play?

The game does not start. Three syntax errors abort loading, leaving a black screen.

PresentationMenus, HUD, graphics, sound

None of it can be seen: without a working build there are no menus, no graphics, no sound.

Code qualityTests, structure, self-corrections

No tests, and three syntax errors that a single start would have shown immediately.

ScopeHow much of the prompt was fulfilled?

A lot is laid out: combat system, reactions, target selection, turn order, sound and post-processing in 6,537 lines. None of it is delivered in the result, because the game does not start.

This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.

What went wrong

  • The delivered game does not start. The model rewrote three loops from forEach to for and each time left the old }); in place: in damage-numbers.js line 142, hud.js line 263 and prompts.js line 274. Each of these spots aborts loading of the whole game, leaving a black screen.
  • With Unsloth's preset for precise coding, presence_penalty 0.0, the model never finished one task: 32,768 tokens over 1,087 seconds, without an answer. Laguna solves the same task in 201 tokens. With presence_penalty 1.5 at temperature 0.6 it reaches the goal with 7,680 tokens in 297 seconds. The first pass of the programming tasks is therefore invalid.
  • Unsloth's example value for the draft depth, 6, is 13 percent below the measured optimum.
  • Same result as Laguna on the programming tasks, but 67 percent slower: the model thinks for 3,000 to 6,600 tokens on every task, even the clearly specified ones, and so produces 2.3 times as many tokens.
  • Not checked is whether MTP holds up under real agent load. With Laguna, draft acceptance collapsed to zero exactly there. The 31.80 t/s come from a single prompt.
  • Not measured yet: speed over context depth and tool calling.
  • On the second image task the model could not read a single price tag. It said so itself and invented nothing.

Synthetic: llama-bench pp512, pp4096 and tg128, plus llama-server for the speculation. Measurement series of 23 July 2026.

Image recognition

A separate test run with the same model on the same machine, not part of the run in which the game was built.

Values read

30 of 50

dashboard, 50 checkable entries

Price tags

0 of 4

read fully correctly

Invented tags

0

none claimed as certain

Stumbling points spotted

1 of 4

places that break the pattern

Provisional: the answer key for the price tags is not fully confirmed yet.

Task 1 · reading values

The test image, a dashboard of my own
The test image. A German-language screenshot of securesight.ai
ZeitIntelligence Index'23'24'25Nov30405060GPT-4GPT-4oClaude 3.5 SonnetClaude Lable 5GPT-5.5Opus 4.6Qwen3.5 27BGemma 4 31BGemma 4 12BMistral Small 4Qwen3.5 8BGemma 3 27BMinistral 3 14BProprietäre FlaggschiffeQwenGemmaMistral
What the model redrew from it as SVG.

What the model read from the tooltip table

Claude Opus 4.5Qwen3.5 27B
Intelligence Index40.837.1
GPQA diamond8784
Humanity's Last Exam2922
Terminal-bench 2.05061

And from the KPI tiles, in German as on the test image

  • QWEN3.5 27B 37.1 · AI Intelligence Index — Open-Weight, kein API-Key, läuft auf der eigenen GPU
  • VS. OPUS 4.5 (49.7) 91 · des Opus-4.5-Niveaus im Gesamtindex. Bei Coding-Benchmarks: gleichauf.
  • PROPRIETÄRE SPITZE · HEUTE 59.9 · Das Rechenzentrum klettert weiter — doch der Abstand schrumpft jedes Quartal.

Task 2 · complex text recognition

What is tested is the reading of tiny, tilted, partly hidden text in a photo with more than forty objects, plus two currency symbols that are hard to tell apart once rotated. The price tags are what the test is run on.

The second test image, a trade fair stand
The second test image. The model named no tag, so nothing is marked in it.

What the model read

No tags named.

The model on the hardest part, quoted in the German it answered in: „Die Preisschilder im unteren linken Bereich (Wertungsbereich: x<700, y>950) sind extrem unscharf, gedreht und teilweise verdeckt. Keine einzige Zahl war sicher lesbar.“

Where the model called itself unsure (3 places)
  • Alle Preisschilder in Bild B - keine Zahlen konnten sicher gelesen werden
  • Genaue Positionsbestimmung welches Schild im Wertungsbereich liegt bei dieser Perspektive
  • Unterscheidung von £ und € bei gedrehten Schildern

The model wrote this list itself, not the scorer, and in German; it is quoted as written. It counts: an admitted “not readable” costs one point, an invented entry costs more.

The complete answer of the model as JSON · The scoring is not published beside it: it names the correct value for every mistake, which would put the answer key in the open.

Try it yourself

  1. Download the weights · Source
  2. Start llama-server with the parameters above
  3. Register the endpoint in VS Code as a custom model and pick the agent

Prompt, agent files and tools: Agent Test Harness

Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run qwen35-122b-a10b-evox2. Measurement data CC-BY-4.0.

Measurement data CC-BY-4.0, code MIT.