Run locally
Laguna S 2.1
AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2
- Hardware
- AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2Strix Halo · Radeon 8060S · gfx1151
- Memory
- 128 GiB unified
- Tasks
- Clair Obscure, Image recognition
- Quantisations
- Q4_K_M
About the model
Mixture-of-Experts model from poolside with 118 billion total parameters and about 8 billion parameters active per token. The model card states 48 layers, of which 12 use global attention and 36 use sliding-window attention with a window size of 512, plus 256 routed experts with top-10 selection and one additional shared expert. Attention is grouped-query with 8 KV heads and head dimension 128, with per-head softplus output gating. Context window 1,048,576 tokens, vocabulary 100,352 tokens, license OpenMDW-1.1.
- Developer
- poolside
- Architecture
- Mixture-of-Experts
- Total parameters
- 118 billion
- Active parameters per token
- about 8 billion
- Layers
- 48 (12 global, 36 sliding window)
- Experts
- 256 routed, top-10, plus 1 shared expert
- Context window
- 1,048,576 tokens
- License
- OpenMDW-1.1
Figures from the vendor: Source
The numbers at a glance
Synthetic lab test with pp512 and tg128, no agent run. The speed therefore comes from a measurement series, not from scored responses of a real run. There is no token balance; it requires many consecutive requests. The curves over context depth come from a separate measurement on the running server with prompts of increasing length. The game next to it comes from a separate run of the same model, built in UD-Q4_K_XL.
Decode without speculation
20.71 t/s
llama-bench tg128, this is how the model runs in operation
Decode at 36,500 tokens
17.8 t/s
Q4_K_M on the running server, at about 60 tokens it is 20.6
Prefill at 512 tokens
338.68 t/s
llama-bench pp512 with --no-mmap
Prefill at 4,096 tokens
275.82 t/s
llama-bench pp4096 with --no-mmap
Programming tasks
5 of 5
with reasoning in 736 seconds, without reasoning only 3 of 5
DFlash in operation
off
measured with Q3_K_M: a net loss in every workload, 0.94 times on code
Memory at 256 K context
89.1 GiB
in use, with KV q8_0 and the draft model for DFlash
Weights
70.01 GiB
Q4_K_M, 117.56 billion parameters, about 8 billion active
Synthetic: llama-bench pp512, pp4096 and tg128, plus llama-server over context depth. Measurement series of 22 and 23 July 2026. · Raw log
Results
Clair Obscur: Expedition 33
Opens in its own layerMouse and keyboard. Enemy attacks are announced, parrying and dodging happen in real time. Needs WebGL.
Loads only on click · 0.3 MB transfer · 1.7 MB unpacked
Measurement series
Decode over context depth
Measured on a running server with real, code-like prompts of increasing length. By 128 K, writing speed drops to less than half.
Prefill over context depth
Up to 32 K, reading the prompt barely loses speed; after that it collapses. Anyone who measures only at shallow depth wrongly takes prefill to be independent of depth.
DFlash over draft depth, measured on a single prompt
The model card recommends a draft depth of 15. On this machine that is two and a half times slower than no speculation at all. Fastest is 3 at 27.9 t/s, but only on this one prompt: in a real agent session acceptance fell to zero, and DFlash is off in operation.
llama.cpp launch parameters
- Context
- 262,144
- KV-Cache
- q8_0 / q8_0
- Micro-Batch
- 512
- Speculative
- draft-dflash · n_max 3 · active in the game run, switched off later
- Build
- 04b2b72
llama-server -m Laguna-S-2.1-UD-Q4_K_XL-00001-of-00003.gguf ^ --host 127.0.0.1 --port 8091 ^ -c 262144 -ngl 999 -fa on ^ -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 ^ --spec-type draft-dflash --spec-draft-n-max 3 ^ --jinja --reasoning on
Assessment
22 / 40
The grade is a personal, and therefore subjective, assessment.
Game feelDoes it run, and does it play?
6/10
The game starts and runs: turn-based combat with four characters against three enemies including a boss, controlled through an action menu with seven actions.
PresentationMenus, HUD, graphics, sound
6/10
Recording available. Its own battle stage with light particles, bars for health and action points on every character.
Code qualityTests, structure, self-corrections
4/10
21 files, 5,994 lines, no tests.
ScopeHow much of the prompt was fulfilled?
6/10
Combat system, reactions, turn order, damage numbers, particles and sound are implemented.
This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.
How do you see it? The grade above is a personal assessment. Here yours counts, independently of it.
What went wrong
- The model card recommends --spec-draft-n-max 15. With Q4_K_M on a single prompt, that gives 8.1 t/s on this machine, two and a half times slower than no speculation at all.
- In a real agent session in VS Code at about 26 K context, draft acceptance fell to zero. Writing dropped to 8.56 t/s, and even requests with 17 new tokens needed almost 12 seconds for the prompt.
- At draft depth 15, which the draft model is trained for, DFlash loses in every workload, measured with Q3_K_M: 0.94 times on code, 0.26 times on prose, 0.38 times on reasoning. DFlash therefore stays off.
- Without reasoning the model is eight times faster, but solves only 3 of 5 programming tasks instead of 5 of 5.
- UD-Q3_K_M writes 51 percent faster, but on two of five tasks it keeps thinking without end and gives no answer even after 24,576 tokens. Q4_K_M completes one of them with 623 tokens. The report therefore rejects Q3.
- In the measurement series over context depth, it took 13.2 minutes to the first token at 128 K.
- The lab values come mostly from Q4_K_M. The DFlash on/off comparison was measured with Q3_K_M, and for the curves over context depth the report names no unambiguous quantization. The model built the game in UD-Q4_K_XL. Its own speed is not measured because there is no server log of that run.
Synthetic: llama-bench pp512, pp4096 and tg128, plus llama-server over context depth. Measurement series of 22 and 23 July 2026.
Image recognition
Laguna S 2.1 is a text-only model and cannot process images. There is therefore no vision test for this model.
Sources
- Measurement repositorygithub.com/KaiFelixBennett/local-ai-amd-benchmark
- Raw log of this rungithub.com/KaiFelixBennett/local-ai-amd-benchmark/blob/main/evidence/reports/laguna-s21-strix-halo-vulkan-benchmark.md
- Source code · Clair Obscure · Q4_K_M · Reasoning on, no limitgithub.com/KaiFelixBennett/local-ai-amd-benchmark/tree/main/benchmarks/laguna-s21/clairobscure
Try it yourself
- Download the weights
- Start llama-server with the parameters above
- Register the endpoint in VS Code as a custom model and pick the agent
Prompt, agent files and tools: Agent Test Harness
Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run laguna-s21-evox2. Measurement data CC-BY-4.0.
Measurement data CC-BY-4.0, code MIT.