Skip to content
benchmark.securesight.ai
DE

Run locally

Qwen3.5-122B-A10B

AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2

Hardware
AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2Strix Halo · Radeon 8060S · gfx1151
Memory
128 GiB unified
Tasks
Labortest, Image recognition
Quantisations
UD-Q4_K_XL

About the model

Sparse Mixture-of-Experts with 122 billion total parameters and 10 billion active per token. The model card states 48 layers in the pattern 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE)), that is three Gated DeltaNet layers per one Gated Attention layer. Each MoE layer holds 256 experts, of which 8 routed plus 1 shared expert are active. The card gives the type as "Causal Language Model with Vision Encoder"; native context length is 262,144 tokens and, per the card, extensible to 1,010,000 tokens.

Typ
Causal Language Model with Vision Encoder (Sparse MoE)
Gesamtparameter
122 B
Aktive Parameter je Token
10 B
Schichten
48
Experten je MoE-Schicht
256
Aktive Experten
8 geroutet + 1 geteilt
Kontextlänge
262 144 Token nativ, erweiterbar bis 1 010 000 Token
Lizenz
apache-2.0

Figures from the vendor: Source · Model card · Vendor

The numbers at a glance

Synthetischer Labortest mit pp512 und tg128, kein Agentenlauf. Die Geschwindigkeit stammt daher aus einer Messreihe, nicht aus gewerteten Antworten eines echten Laufs. Eine Tokenbilanz und eine Kurve über die Kontexttiefe gibt es nicht, beides setzt viele aufeinanderfolgende Anfragen voraus. Das Artefakt daneben stammt aus einem eigenen Durchgang desselben Modells.

Decode mit MTP

31.80 t/s

Entwurfstiefe 2, ein einzelner Prompt, das 1,55fache des Werts ohne Spekulation

Decode bei den Aufgaben

26.5 bis 28,0 t/s

mit MTP, Entwurfstiefe 2, über die fünf Programmieraufgaben

Decode ohne Spekulation

20.52 t/s

llama-bench tg128

Annahme der Entwürfe

0.866

bei Entwurfstiefe 2, im Mittel 2,73 Token je Schritt

Prefill bei 512 Token

245.71 t/s

llama-bench pp512

Prefill bei 4 096 Token

232.68 t/s

nur 5 Prozent unter dem Wert bei 512 Token

Programmieraufgaben

5 von 5

gegen versteckte Tests ausgeführt, 1 232 Sekunden gesamt

Gewichte

73.23 GiB

UD-Q4_K_XL, 124,64 Mrd. Parameter, rund 10 Mrd. aktiv

Synthetisch: llama-bench pp512, pp4096 und tg128, dazu llama-server für die Spekulation. Messreihe vom 23.07.2026. · Raw log

Results

Das Spiel, das dieser Lauf abgeliefert hat, startet nicht: drei Syntaxfehler brechen das Laden ab, übrig bleibt ein schwarzer Bildschirm. Deshalb gibt es hier weder eine Aufnahme noch eine Version zum Ausprobieren. Der Quelltext liegt so, wie das Modell ihn abgegeben hat, im Belegrepo.

Measurement series

Geschwindigkeit über die Entwurfstiefe

MTP sagt mehrere Token auf einmal voraus, und das Modell prüft sie in einem Schritt. Wie viele es je Schritt versucht, stellt --spec-draft-n-max ein. Mehr ist nicht automatisch schneller: bei 6 wird so viel verworfen, dass der Gewinn schrumpft.

192633 0 · 20.55 t/s · ohne Spekulation 2 · 31.80 t/s · Annahme 0,866, Länge 2,73 3 · 31.05 t/s · Annahme 0,726, Länge 3,18 4 · 31.89 t/s · Annahme 0,725, Länge 3,90 6 · 27.68 t/s · Annahme 0,530, Länge 4,18 im BetriebBeispielwert von Unsloth 02346 t/s
Entwurfstiefe, --spec-draft-n-max · Ein Durchlauf je Tiefe, frisch geladener Server, derselbe Prompt, KV q8_0, Denken an. 2 und 4 liegen 0,3 Prozent auseinander und gelten als gleich schnell. Gewählt ist 2, weil die Annahme mit 0,866 mehr Reserve lässt. Bericht, Abschnitt 5.3.

llama.cpp launch parameters

Context
262,144
KV-Cache
q8_0 / q8_0
Micro-Batch
512
Speculative
draft-mtp · n_max 2 · Annahme 0,866
Build
04b2b72
llama-server -m Qwen3.5-122B-A10B-UD-Q4_K_XL-00001-of-00003.gguf ^
  --mmproj mmproj-F16.gguf --host 127.0.0.1 --port 8098 ^
  -c 262144 -ngl 999 -fa on ^
  -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 ^
  --spec-type draft-mtp --spec-draft-n-max 2 ^
  --temp 0.6 --top-p 0.95 --presence-penalty 1.5 ^
  --jinja --chat-template-file qwen-fixed-v21.3.jinja ^
  --reasoning on --reasoning-format deepseek

Assessment

4 / 40

The grade is a personal, and therefore subjective, assessment.

Game feelDoes it run, and does it play?

Das Spiel startet nicht. Drei Syntaxfehler brechen das Laden ab, es bleibt ein schwarzer Bildschirm.

PresentationMenus, HUD, graphics, sound

Davon ist nichts zu sehen: ohne lauffähigen Build keine Menüs, keine Grafik, kein Ton.

Code qualityTests, structure, self-corrections

Keine Tests, und drei Syntaxfehler, die ein einziger Start sofort gezeigt hätte.

ScopeHow much of the prompt was fulfilled?

Angelegt ist viel: Kampfsystem, Reaktionen, Zielauswahl, Zugreihenfolge, Ton und Nachbearbeitung in 6 537 Zeilen. Erfüllt ist davon im Ergebnis nichts, weil das Spiel nicht startet.

This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.

What went wrong

  • Das ausgelieferte Spiel startet nicht. Das Modell hat drei Schleifen von forEach auf for umgeschrieben und dabei jedes Mal das alte }); stehen lassen: in damage-numbers.js Zeile 142, hud.js Zeile 263 und prompts.js Zeile 274. Jede dieser Stellen bricht das Laden des ganzen Spiels ab, übrig bleibt ein schwarzer Bildschirm.
  • Mit Unsloths Voreinstellung für präzises Programmieren, presence_penalty 0,0, hat das Modell eine Aufgabe nie beendet: 32 768 Token über 1 087 Sekunden, ohne Antwort. Laguna löst dieselbe Aufgabe in 201 Token. Mit presence_penalty 1,5 bei Temperatur 0,6 kommt es mit 7 680 Token in 297 Sekunden zum Ziel. Der erste Durchgang der Programmieraufgaben ist deshalb ungültig.
  • Unsloths Beispielwert für die Entwurfstiefe, 6, liegt 13 Prozent unter dem gemessenen Optimum.
  • Bei den Programmieraufgaben dasselbe Ergebnis wie Laguna, aber 67 Prozent langsamer: das Modell denkt bei jeder Aufgabe 3 000 bis 6 600 Token nach, auch bei den eindeutig beschriebenen, und erzeugt so 2,3 mal so viele Token.
  • Nicht geprüft ist, ob MTP unter echter Agentenlast hält. Bei Laguna brach die Annahme der Entwürfe genau dort auf null ein. Die 31,80 t/s stammen aus einem einzelnen Prompt.
  • Noch nicht gemessen: die Geschwindigkeit über die Kontexttiefe und Tool Calling.
  • Bei der zweiten Bildaufgabe hat das Modell kein einziges Preisschild lesen können. Es hat das selbst gesagt und nichts erfunden.

Synthetisch: llama-bench pp512, pp4096 und tg128, dazu llama-server für die Spekulation. Messreihe vom 23.07.2026.

Image recognition

Eigener Testlauf mit demselben Modell auf derselben Maschine, nicht Teil des Laufs, in dem das Spiel entstand.

Values read

30 of 50

dashboard, 50 checkable entries

Price tags

0 of 4

read fully correctly

Invented tags

0

none claimed as certain

Traps passed

1 of 4

places that break the pattern

Provisional: the answer key for the price tags is not fully confirmed yet.

Task 1 · reading values

The test image, a dashboard of my own
The test image. A screenshot of securesight.ai
ZeitIntelligence Index'23'24'25Nov30405060GPT-4GPT-4oClaude 3.5 SonnetClaude Lable 5GPT-5.5Opus 4.6Qwen3.5 27BGemma 4 31BGemma 4 12BMistral Small 4Qwen3.5 8BGemma 3 27BMinistral 3 14BProprietäre FlaggschiffeQwenGemmaMistral
What the model redrew from it as SVG.

What the model read from the tooltip table

Claude Opus 4.5Qwen3.5 27B
Intelligence Index40.837.1
GPQA diamond8784
Humanity's Last Exam2922
Terminal-bench 2.05061

And from the KPI tiles

  • QWEN3.5 27B 37.1 · AI Intelligence Index — Open-Weight, kein API-Key, läuft auf der eigenen GPU
  • VS. OPUS 4.5 (49.7) 91 · des Opus-4.5-Niveaus im Gesamtindex. Bei Coding-Benchmarks: gleichauf.
  • PROPRIETÄRE SPITZE · HEUTE 59.9 · Das Rechenzentrum klettert weiter — doch der Abstand schrumpft jedes Quartal.

Task 2 · complex text recognition

What is tested is the reading of tiny, tilted, partly hidden text in a photo with more than forty objects, plus two currency symbols that are hard to tell apart once rotated. The price tags are what the test is run on.

The second test image, a trade fair stand
The second test image. The model named no tag, so nothing is marked in it.

What the model read

No tags named.

The model on the hardest part: „Die Preisschilder im unteren linken Bereich (Wertungsbereich: x<700, y>950) sind extrem unscharf, gedreht und teilweise verdeckt. Keine einzige Zahl war sicher lesbar.“

Where the model called itself unsure (3 places)
  • Alle Preisschilder in Bild B - keine Zahlen konnten sicher gelesen werden
  • Genaue Positionsbestimmung welches Schild im Wertungsbereich liegt bei dieser Perspektive
  • Unterscheidung von £ und € bei gedrehten Schildern

The model wrote this list itself, not the scorer. It counts: an admitted “not readable” costs one point, an invented entry costs more.

The complete answer of the model as JSON · The scoring is not published beside it: it names the correct value for every mistake, which would put the answer key in the open.

Try it yourself

  1. Download the weights · Source
  2. Start llama-server with the parameters above
  3. Register the endpoint in VS Code as a custom model and pick the agent

Prompt, agent files and tools: Agent Test Harness

Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run qwen35-122b-a10b-evox2. Measurement data CC-BY-4.0.

Measurement data CC-BY-4.0, code MIT.