Skip to content

Run locally

Laguna S 2.1

AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2

Hardware
AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2Strix Halo · Radeon 8060S · gfx1151
Memory
128 GiB unified
Tasks
Clair Obscure, Image recognition
Quantisations
Q4_K_M

About the model

Mixture-of-Experts model from poolside with 118 billion total parameters and about 8 billion parameters active per token. The model card states 48 layers, of which 12 use global attention and 36 use sliding-window attention with a window size of 512, plus 256 routed experts with top-10 selection and one additional shared expert. Attention is grouped-query with 8 KV heads and head dimension 128, with per-head softplus output gating. Context window 1,048,576 tokens, vocabulary 100,352 tokens, license OpenMDW-1.1.

Developer
poolside
Architecture
Mixture-of-Experts
Total parameters
118 billion
Active parameters per token
about 8 billion
Layers
48 (12 global, 36 sliding window)
Experts
256 routed, top-10, plus 1 shared expert
Context window
1,048,576 tokens
License
OpenMDW-1.1

Figures from the vendor: Source

The numbers at a glance

Synthetic lab test with pp512 and tg128, no agent run. The speed therefore comes from a measurement series, not from scored responses of a real run. There is no token balance; it requires many consecutive requests. The curves over context depth come from a separate measurement on the running server with prompts of increasing length. The game next to it comes from a separate run of the same model, built in UD-Q4_K_XL.

Decode without speculation

20.71 t/s

llama-bench tg128, this is how the model runs in operation

Decode at 36,500 tokens

17.8 t/s

Q4_K_M on the running server, at about 60 tokens it is 20.6

Prefill at 512 tokens

338.68 t/s

llama-bench pp512 with --no-mmap

Prefill at 4,096 tokens

275.82 t/s

llama-bench pp4096 with --no-mmap

Programming tasks

5 of 5

with reasoning in 736 seconds, without reasoning only 3 of 5

DFlash in operation

off

measured with Q3_K_M: a net loss in every workload, 0.94 times on code

Memory at 256 K context

89.1 GiB

in use, with KV q8_0 and the draft model for DFlash

Weights

70.01 GiB

Q4_K_M, 117.56 billion parameters, about 8 billion active

Synthetic: llama-bench pp512, pp4096 and tg128, plus llama-server over context depth. Measurement series of 22 and 23 July 2026. · Raw log

Results

Clair Obscur: Expedition 33, built by Laguna S 2.1. A battle of four characters against three enemies including a boss, controlled through an action menu, uncut.

Clair Obscur: Expedition 33

Opens in its own layer

Mouse and keyboard. Enemy attacks are announced, parrying and dodging happen in real time. Needs WebGL.

Loads only on click · 0.3 MB transfer · 1.7 MB unpacked

Measurement series

Decode over context depth

Measured on a running server with real, code-like prompts of increasing length. By 128 K, writing speed drops to less than half.

112132 4.3 K · 29.28 t/s · Prompt with 4,267 tokens 33.5 K · 18.13 t/s · Prompt with 33,477 tokens 133.8 K · 13.28 t/s · Prompt with 133,839 tokens 4.3 K33.5 K133.8 K t/s
Context depth in thousands of tokens · llama-server, Vulkan, KV q8_0, without speculation. Section 5.7 does not say which quantization was running. Section 6.2c lists the 18.13 t/s at 32 K as Q3 and later, with a newer build, measures 26.6 t/s with Q3_K_M and 17.8 t/s with Q4_K_M at 36,500 tokens. According to section 5.7.2, the value at 4 K is not directly comparable with the 20.71 t/s from llama-bench: that used KV f16 and a different timing method.

Prefill over context depth

Up to 32 K, reading the prompt barely loses speed; after that it collapses. Anyone who measures only at shallow depth wrongly takes prefill to be independent of depth.

109190271 4.3 K · 252.90 t/s · 16.9 seconds to the first token 33.5 K · 239.80 t/s · 121.9 seconds to the first token 133.8 K · 126.40 t/s · 794 seconds, 13.2 minutes to the first token 4.3 K33.5 K133.8 K t/s
Context depth in thousands of tokens · From 4 K to 32 K 5 percent less, from 32 K to 128 K a further 47 percent. The same measurement series as above, here too without the quantization being stated. Report, section 5.7.

DFlash over draft depth, measured on a single prompt

The model card recommends a draft depth of 15. On this machine that is two and a half times slower than no speculation at all. Fastest is 3 at 27.9 t/s, but only on this one prompt: in a real agent session acceptance fell to zero, and DFlash is off in operation.

51831 without speculation 20.4 t/s 2 · 26.00 t/s · acceptance 0.645 3 · 27.90 t/s · acceptance 0.559 4 · 24.70 t/s · acceptance 0.443 5 · 23.10 t/s · acceptance 0.391 6 · 19.60 t/s · acceptance 0.293 8 · 9.70 t/s · acceptance 0.265 15 · 8.10 t/s · acceptance 0.150 fastest value, one promptmodel card 23456815 t/s
Draft depth, --spec-draft-n-max · Q4_K_M, Poolside fork 04b2b72, temperature 0, the same prompt. The line without speculation comes from the official llama.cpp, not from the fork, report section 9. That DFlash loses in operation is stated in sections 5.8.1 and 6.2b of the report.

llama.cpp launch parameters

Context
262,144
KV-Cache
q8_0 / q8_0
Micro-Batch
512
Speculative
draft-dflash · n_max 3 · active in the game run, switched off later
Build
04b2b72
llama-server -m Laguna-S-2.1-UD-Q4_K_XL-00001-of-00003.gguf ^
  --host 127.0.0.1 --port 8091 ^
  -c 262144 -ngl 999 -fa on ^
  -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 ^
  --spec-type draft-dflash --spec-draft-n-max 3 ^
  --jinja --reasoning on

Assessment

22 / 40

The grade is a personal, and therefore subjective, assessment.

Game feelDoes it run, and does it play?

The game starts and runs: turn-based combat with four characters against three enemies including a boss, controlled through an action menu with seven actions.

PresentationMenus, HUD, graphics, sound

Recording available. Its own battle stage with light particles, bars for health and action points on every character.

Code qualityTests, structure, self-corrections

21 files, 5,994 lines, no tests.

ScopeHow much of the prompt was fulfilled?

Combat system, reactions, turn order, damage numbers, particles and sound are implemented.

This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.

What went wrong

  • The model card recommends --spec-draft-n-max 15. With Q4_K_M on a single prompt, that gives 8.1 t/s on this machine, two and a half times slower than no speculation at all.
  • In a real agent session in VS Code at about 26 K context, draft acceptance fell to zero. Writing dropped to 8.56 t/s, and even requests with 17 new tokens needed almost 12 seconds for the prompt.
  • At draft depth 15, which the draft model is trained for, DFlash loses in every workload, measured with Q3_K_M: 0.94 times on code, 0.26 times on prose, 0.38 times on reasoning. DFlash therefore stays off.
  • Without reasoning the model is eight times faster, but solves only 3 of 5 programming tasks instead of 5 of 5.
  • UD-Q3_K_M writes 51 percent faster, but on two of five tasks it keeps thinking without end and gives no answer even after 24,576 tokens. Q4_K_M completes one of them with 623 tokens. The report therefore rejects Q3.
  • In the measurement series over context depth, it took 13.2 minutes to the first token at 128 K.
  • The lab values come mostly from Q4_K_M. The DFlash on/off comparison was measured with Q3_K_M, and for the curves over context depth the report names no unambiguous quantization. The model built the game in UD-Q4_K_XL. Its own speed is not measured because there is no server log of that run.

Synthetic: llama-bench pp512, pp4096 and tg128, plus llama-server over context depth. Measurement series of 22 and 23 July 2026.

Image recognition

Laguna S 2.1 is a text-only model and cannot process images. There is therefore no vision test for this model.

Try it yourself

  1. Download the weights
  2. Start llama-server with the parameters above
  3. Register the endpoint in VS Code as a custom model and pick the agent

Prompt, agent files and tools: Agent Test Harness

Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run laguna-s21-evox2. Measurement data CC-BY-4.0.

Measurement data CC-BY-4.0, code MIT.