Skip to content

Tests and first-hand experience

Local AI on affordable hardware

Which models deliver which results, and how fast? No standard benchmarks, but image recognition and agentic coding with VS Code tried out first hand on local, affordable AMD hardware.

8runs with measurements
2machines
37.9 hserver runtime

See every run

Human in the loop

Kai Bennett

I am Kai Bennett, a software architect, and I have been enthusiastic about local AI since September 2025. Here I share my findings and the results I measured on the two devices.

Certified Software Architect (iSAQB CPSA-F/A) · Co-founder of securesight.ai

LinkedInGitHubXingEmail

How this is measured

Personal recommendation

My pick per machine.

With these two models I ran the longest agentic autonomous jobs, of up to 13.6 h, and at the same time got very good results at a usable median speed of over 20 t/s. For me, on this hardware and as of September 2026, that is the best compromise between quality and speed, without paying for tokens, without depending on cloud providers and without sharing my data.

AI MAX 395

Qwen3.8-Flash-Next

UD-Q4_K_XL

21.76t/s Median

R9700

Qwen3.8-27B

UD-Q4_K_XL

33.69t/s Median

Why these two

AMD AI MAX 395 · AMD Halo · 13.6 h

Single shot

Qwen3.8-Flash-Next builds a finished game of — lines of code including tests, autonomously, from a single prompt.

13.6 hruntime
459requests
787ktokens generated

Play it now

Radeon AI PRO R9700

The fastest measured run.

Qwen3.8-27B UD-Q4_K_XL at over 30 t/s on average, and still around 30 at long context. That is a genuinely usable speed, and this is a dense model in which all 27 billion parameters are active at once.

33.69t/s median
45.34t/s p90
9.0 hruntime

Every speed measured

Machine 1

Radeon AI PRO R9700

gfx1201 · RDNA4 · 32 GB dedicated, run headless.
Intel Core Ultra 7 265KF · 63,6 GiB RAM · ASUS PRIME Z890-P.

Windows 11 Pro 26200 · Adrenalin 26.6.4 · amdvlk64 9.2.10.395

Machine 2

AMD AI MAX 395 (AMD Halo)

Strix Halo · Radeon 8060S · gfx1151 · 128 GiB unified LPDDR5X-8533, around 256 GB/s theoretical bandwidth.

Windows 11 Pro 26200 · Driver 32.0.31021.5001 · Vulkan heap 98.123 MiB

Software

llama.cpp over Vulkan

Every run on llama.cpp with the Vulkan backend, no ROCm, no BIOS changes. Served on the local network and wired into VS Code Copilot Chat as a custom endpoint.

Builds b10717 · bd9bd1b · b9985 · 580e88d · poolside 04b2b72

Personal recommendation

These two models gave me the longest autonomous agent jobs, up to 13.6 h, while still producing very good results at a usable median speed of over 20 t/s. On this hardware, as of September 2026, that is the best compromise I have found.

AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2

Qwen3.8-Flash-Next

UD-Q4_K_XL · MTP, Shared-Q8_0, n-max 2

t/s median
21.76
Runtime
13.6 h
Requests
459

Not the fastest and not the highest quality either, but it makes use of the advantages of the AMD Strix Halo, meaning it uses the entire VRAM, already includes the Qwen 4 architecture and delivers well at a steady speed.

Raw log

Radeon AI PRO R9700

Qwen3.8-27B

UD-Q4_K_XL · draft-mtp + ngram-mod 24/48/64

t/s median
33.69
Runtime
9.0 h
Requests
108

The first model I would also recommend for coding on a GPU with 24–32 GB of VRAM. In Q4 it runs stably and agentically for hours, writes tests and reviews itself. If I were to recommend one model in the range up to 32 GB of VRAM, it would definitely be this one.

Raw log

The models

Videos, image recognition and measurements on the detail pages

For example a test inspired by the retro game “Moorhuhn”, created from just one prompt, playable for yourself. Every measurement including the speed and the tokens consumed. Image recognition is tested too, an important ability for the models so they can look at their own results and improve them.

For this, an agent was configured in Visual Studio Code carrying the corresponding instructions.

14 of 14 detail pages built·0 to come

/

Type to filter · ↑ ↓ to select · Enter opens the detail page

Qwen3.8-Flash-NextMoorhuhn · Clair Obscure UD-Q4_K_XL 20.37 t/s median 33.3M tokens 1185minutes AI MAX 395 Raw log +1 Details
Qwen3.8-27BMoorhuhn · Clair Obscure UD-Q4_K_XL · UD-Q6_K_M 31.77 t/s median 11.7M tokens 1043minutes R9700 Raw log +1 Details
Qwen3.6-27BMoorhuhn · Clair Obscure UD-Q6_K_XL 33.45 t/s median 28.9M tokens 153minutes R9700 Raw log Details
Qwen3.5-122B-A10BLabortest UD-Q4_K_XL 31.80 t/s median M tokens minutes AI MAX 395 Raw log Details
Laguna S 2.1Clair Obscure Q4_K_M 20.71 t/s median M tokens minutes AI MAX 395 Raw log Details
DeepSeek-V4-Flash-0731Clair Obscure UD-IQ3_XXS 7.02 t/s median 8.0M tokens 502minutes AI MAX 395 Raw log Details
Opus 5 ULTRACODEClair Obscure · Moorhuhn t/s median M tokens minutes Cloud not measured Details
GPT 5.6 SolMoorhuhn t/s median M tokens minutes Cloud not measured Details
Sonnet 5Moorhuhn t/s median M tokens 70minutes Cloud not measured Details

Each model appears once, with every task it worked on. Where several runs exist, the median is computed across all scored responses of those runs together, not averaged from the individual medians; tokens and minutes are sums. The per-run figures are on the detail page.
M tokens used is what a paid provider would have charged for: input plus output, with every single request counted at its full context, not just at what was new. An agent run resends the same history hundreds of times, which is why input runs a multiple above output. Short responses are included, even the ones that do not enter the speed figure. The split into input and output is in each row's tooltip.

The best result so far

A 13.5 h run builds a complete small game, including settings and saved progress among other things.

Qwen3.8-Flash-Next on the AMD AI MAX 395, 177 billion parameters, roughly six of them active. 459 requests, 786,506 tokens generated, not a single intervention. Graphics, menus, sound, backend: all of it generated along the way. The large models can do this too, of course, but in places I found the result even better than Sonnet 5's, and it runs on hardware costing around 2,600 € (a Bosgame M5, for instance).

Game scene: moorland at dusk, a scarecrow and tin cans as targets, with the timer and score displayed at the top.

Built by Qwen3.8-Flash-Next · AMD AI MAX 395

13 h 35on one prompt
459requests
786,506tokens generated

Read the prompt Raw log

The result

Moorhuhn & Clair Obscure single-shot benchmark

Recordings from the delivered builds: menus, game modes, scoring, unlocks.

Moorland Mayhem — Featherstorm · built by Qwen3.8-27B UD-Q4_K_XL on the Radeon AI PRO R9700 · from a single prompt. Settings, statistics, a full round and the scoreboard. Uncut, 1 minute 44.
Qwen3.6-27B Q6Clair Obscure · combat, own recording
Qwen3.6-27B Q6Moorhuhn · 1,125 points
Laguna S 2.1 · AMD AI MAX 395Clair Obscure · turn-based combat
Clair Obscure — Expedition 33 · Qwen3.8-27B UD-Q6 on the Radeon AI PRO R9700 · night scene with dialogue, portal, combat arena with targeting, ability menu and title screen. Uncut, 3 minutes 41.

Reference run

Same task, different model.

Qwen3.8-Flash-Next on the AMD AI MAX 395 (AMD Halo), 177 billion parameters in total, roughly six of them active. Combat arena with magic circle, action menu, status bars and spell effects.

Clair Obscure — Expedition 33 · Qwen3.8-Flash-Next UD-Q4_K_XL on the AMD AI MAX 395 (AMD Halo) · uncut, 3 minutes 16.

Agentic development

Single shot. 13.5 h.

The model runs on the local network and is registered in VS Code Copilot Chat as a custom endpoint, exactly where most developers already are. With tool use it writes tests, creates files and folders, and produces screenshots and Playwright tests to check its own work.

The run is finished: 13 hours 36 minutes on a single prompt, 459 requests, 786,506 tokens generated and 1,051,292 read. Counting cached context, the model read 24.9 million tokens. The context was never truncated once. The peak was 98,109 of 131,072.

The screenshot shows an intermediate state after 8 hours 47 minutes: the seventh of thirteen work packages, 40 files changed, the context window at 51 per cent.

VS Code with a local model: llama.cpp console and the agent at work
Qwen3.8-Flash-Next Q4XL + MTP · 177B/6B active · 131k · Vision · :8099
Server console on top, the agent below: 40 files changed (+6787/−40), context window 67.3 K of 131 K.
1,835,623tokens generated
869measured responses
42 hruntime

Vision

Image recognition

Many local models can read images. Testing uses my own photographs, to find out how well image recognition works.

Dashboard screenshot used as a test image
Reading valuesA dashboard of my own: tooltip table, KPI tiles, active filters, footnotes. The model returns them as JSON and redraws the chart as SVG.
A dense photograph used as a test image
Reading in clutterA trade-fair stand with more than forty objects. The count is how many price tags are read correctly, and how many are invented.

What the models recognised

In each image exactly one place breaks the pattern. That is a challenge for the models, one where they have to make out detail. Those four places are therefore reported separately.

What is tested
Qwen3.8-27BUD-Q4_K_XL · Radeon AI PRO R9700
Qwen3.6-27BUD-Q6_K_XL · Radeon AI PRO R9700
Qwen3.5-122B-A10BUD-Q4_K_XL · AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2
Sonnet 5— · Cloud reference
Values read correctly 44 of 5050 of 5030 of 5046 of 50
Row Terminal-Bench 2.01 recognisedrecognisedrecognisedrecognised
Tile value 40.72 recognisedrecognisednot recognisednot recognised
Price tag JOKER3 not recognisedrecognisednot recognisedrecognised
White mask, unnamed4 not recognisedrecognisednot recognisedorder spotted, amounts swapped
Price tags fully correct5 2 of 44 of 40 of 41 of 4
Invented tags6 4 · 1 claimed as certain005 · none claimed as certain
  1. Row Terminal-Bench 2.0 In every other row of the table the larger value is on the left. In this one it is not. A model that skims the table and continues the pattern flips it.
  2. Tile value 40.7 The tile gives 40.7 for the same model, the tooltip gives 40.8. Both numbers really are in the image. Harmonising them means smoothing instead of reading.
  3. Price tag JOKER On every other tag the pound price is higher than the euro price. On this one it is the other way round.
  4. White mask, unnamed On every other tag the pound is on top and the euro below. On this one the order is reversed.
  5. Price tags fully correct Scored only inside the defined image area, where the tags face the camera. Correct means: name matched in substance, both amounts exact, currency order right. Two out of three is not enough.
  6. Invented tags Tags named that do not appear in the answer key. Reported separately: how many of them were claimed as certain. An admitted “not readable” costs one point, an invented entry costs more.

Provisional. The answer key for the price tags is confirmed for four of seven tags. Precision and recall are computed against those four only and will change once the rest is confirmed. The four traps are unaffected, they are confirmed.

Speed

33.7t/s

This is how fast the fastest run writes.

Median over 82 responses. What counts is not the peak speed reached, but the speed you get on average while working through a task.

Hardware

Context depth

Benchmark speeds distort the picture.

Because they almost always measure at short context. In real use the context grows continuously and speed drops steadily. That is why this page shows median speeds and how they develop with context size.

Prefill · Qwen3.8-27B UD-Q6 · R9700 Report

Decode · Qwen3.8-Flash-Next · AMD AI MAX 395 (AMD Halo) Raw log

On the R9700 the prefill rate drops by a factor of 4.2 between 8 K and 164 K. On the right, the decode rate of Qwen3.8-Flash-Next from the Moorhuhn run: 346 responses of 200 tokens or more, each with its context depth from the log, in six equally sized groups, measured between 34 K and 71 K of context, the range that matters in everyday use. Across that span it drops by a factor of 1.18. The model ran especially steadily and reliably.

Fine-tuning

Its all about settings

The parameters in the model cards are not always optimal (for AMD). These three, for example, I changed.

spec-draft-n-max · Laguna S 2.1

15 → 3

Lowering “spec-draft-n-max” produces more throughput

With the documented --spec-draft-n-max 15 decode falls to 8.1 t/s, below the 20.55 achieved without any speculation at all. With 3 it rises to 27.9.

Series

Sampling · Qwen3.5-122B

0.0 → 1.5

Without the setting that lowers the likelihood of repetition, the answer never ends

With presence_penalty 0.0 one task ran to the end of the context: 32,768 tokens, never finishing. The same task needs 201 tokens with the setting switched on.

Series

Memory · Strix Halo

25.6 GiB

Raising the reservation did not help

A larger UMA reservation brought no gain and would have taken memory away from the operating system. Even at 262,144 context, 25.6 GiB remained free.

Measurement

Model cards

19 launch configurations.

Every row is a script that started exactly one such server. Each set of parameters can be copied on its own, complete, with every flag.

ModelHardware Context-ub KV k / vSpeculation Vision
Qwen3.6-27B-UD-Q6_K_XLstart_qwen36_27b_q6_coding.bat R9700 163,840 512 q8_0 / q4_0 Vision
Qwen3.6-27B-UD-Q6_K_XL-mtpstart_qwen36_27b_q6_mtp_coding.bat R9700 172,032 512 q8_0 / q4_0 draft-mtp Vision
start_qwen36_35b_a3b_mtp.bat R9700 172,032 — / —
start_qwen38_27b_coding.bat R9700 262,144 — / —
start_qwen38_27b_ngram_test.bat R9700 262,144 — / —
start_qwen38_27b_q4xl_coding.bat R9700 262,144 — / —
start_qwen38_27b_q4xl_medium.bat R9700 262,144 — / —
Qwen3.8-27B-UD-Q6_K_Mstart_qwen38_27b_test_rmkq1_ub128.bat R9700 262,144 128 q8_0 / turbo4 draft-mtp Vision
Qwen3.8-27B-UD-Q6_K_Mstart_qwen38_27b_test_ub256.bat R9700 262,144 256 q8_0 / turbo4 draft-mtp Vision
Qwen3.8-27B-UD-Q6_K_Mstart_qwen38_27b_test_ub64.bat R9700 262,144 64 q8_0 / turbo4 draft-mtp Vision
DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004start-deepseek-v4-flash-0731-DSPARK.bat AI MAX 395 131,072 512 f16 / f16 draft-dspark
Qwen3.6-35B-A3B-UD-Q4_K_XLstart-qwen-3.6-35b-a3b-vision.bat AI MAX 395 262,144 — / — Vision
Qwen3.5-122B-A10B-UD-Q4_K_XL-00001-of-00003start-qwen35-122b-A10B-MTP-NOTHINK.bat AI MAX 395 262,144 512 q8_0 / q8_0 draft-mtp Vision
Qwen3.5-122B-A10B-UD-Q4_K_XL-00001-of-00003start-qwen35-122b-A10B-MTP.bat AI MAX 395 262,144 512 q8_0 / q8_0 draft-mtp Vision
Qwen3.6-27B-UD-Q6_K_XLstart-qwen36-27b-MTP.bat AI MAX 395 262,144 512 q8_0 / q4_0 draft-mtp Vision
Qwen3.8-27B-UD-Q6_K_XLstart-qwen38-27b-FALLBACK-8090.bat AI MAX 395 32,768 512 q8_0 / q4_0 draft-mtp Vision
Qwen3.8-27B-UD-Q6_K_XL-240819start-qwen38-27b-MTP-240819.bat AI MAX 395 262,144 512 q8_0 / q4_0 draft-mtp Vision
Qwen3.8-27B-UD-Q6_K_XLstart-qwen38-27b-MTP.bat AI MAX 395 262,144 512 q8_0 / q4_0 draft-mtp Vision
Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004start-qwen38-flash-next.bat AI MAX 395 131,072 512 q8_0 / q8_0 draft-mtp Vision

Agent Test Harness

Everything you need is right here.

Two different things are measured. The run shows speed, stamina and whether something runnable exists at the end. The detail page shows: design ability, image recognition, and whether it holds to a contract without inventing.

01

The prompts

Two prompts, from which every result on this page came. Unchanged, exactly as the models received them. They are deliberately built as opposites.

Moorhuhn
21 KB · breadth and stamina. Explicitly demands a finished game with modes, menus, statistics and tests, not a scaffold.
Clair Obscure
15 KB · depth instead of breadth. A single fight, but with parry and dodge on tight timing windows. The prompt itself calls that layer the hardest part.

One prompt can be passed by diligence, the other cannot.

02

The agents

One agent for VS Code Copilot Chat. Drop it into .github/agents/ and it shows up in the chat picker.

bilderkennung.agent.md GitHub
The vision test. It deliberately has no terminal, so the model has to look at the reference images instead of cropping them with an image library.

llama-server speaks the OpenAI API. VS Code accepts such endpoints through “bring your own key”, without a Copilot subscription and without signing in. The launch parameters for each run are given under model cards.

VS Code documentation Harness in the repository

03

The detail pages

Every run has its own detail page with the same eleven sections: architecture, measurements, artefact, image recognition, speed, configuration, quality, errors, sources, reproduction.

The eleven sections The design rules

Test bench

Two machines. Both under 3,500 euros.

Copy the launch parameters of the fastest measured run on each machine.

Radeon AI PRO R9700

gfx1201 · RDNA4 · 32 GB dedicated

from ~1,500 €

Allocatable
32,624 MiB
CPU
Intel Core Ultra 7 265KF
RAM
63.6 GiB
Motherboard
ASUS PRIME Z890-P WIFI, BIOS 2401
Driver
Adrenalin 26.6.4
Vulkan ICD
amdvlk64 9.2.10.395

Fastest run: Qwen3.8-27B UD-Q4_K_XL · 33.69 t/s

llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf ^ --load-mode none --mmproj mmproj-F16.gguf --image-min-tokens 1024 ^ --alias Qwen3.8-27B -dev Vulkan0 --host 127.0.0.1 --port 8080 ^ -c 262144 -ngl 99 -fa on --fit off -ot "token_embd\.weight=Vulkan0" ^ -ctk q8_0 -ctv q8_0 -b 2048 -ub 288 ^ --parallel 1 --kv-unified ^ --ctx-checkpoints 4 --checkpoint-min-step 16384 --cache-ram 0 ^ --spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 ^ --spec-ngram-mod-n-match 24 --spec-ngram-mod-n-min 48 --spec-ngram-mod-n-max 64 ^ --spec-draft-device Vulkan0 --spec-draft-ngl 99 ^ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 ^ --presence-penalty 0.0 --repeat-penalty 1.0 ^ --jinja --reasoning on --reasoning-effort xhigh --reasoning-format deepseek --reasoning-budget -1

AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2

Strix Halo · Radeon 8060S · gfx1151 · 128 GiB unified

from ~1,800 €

Memory
128 GiB LPDDR5X-8533, 8 channels
Bandwidth
~256 GB/s theoretical
Visible to Windows
63.6 GiB
UMA reservation
64.4 GiB
Vulkan heap
98,123 MiB / 93,217 MiB free
Driver
32.0.31021.5001

Fastest run: Qwen3.8-Flash-Next UD-Q4_K_XL · 21.76 t/s

llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf ^ --mmproj mmproj-F16.gguf --alias qwen3.8-flash-next ^ --host 127.0.0.1 --port 8099 --device Vulkan0 ^ --gpu-layers all --n-cpu-moe 0 --fit off -fa on ^ --load-mode mmap --lazy-mode on ^ --ctx-size 262144 --parallel 1 --kv-unified ^ -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 ^ --ctx-checkpoints 4 --checkpoint-min-step 4096 ^ --jinja --reasoning on --reasoning-format deepseek --reasoning-effort xhigh

Three numbers for the same memory. The AMD AI MAX 395 has 128 GiB of unified LPDDR5X, not 128 GB of VRAM. Windows sees 63.6 GiB, the BIOS reservation is 64.4 GiB, and Vulkan reports 98,123 MiB.

Verifiability

Trust, but verify

Logs and measurements are in the public repository and were taken on the two systems described above. Depending on your exact hardware and software setup, your numbers may of course differ somewhat.

1

Raw data in the public repository

The llama.cpp server logs and the measurement tables sit unchanged in evidence/, around 3.8 MB in total, with checksums. Every model row and every table line links exactly the file its numbers come from.

2

Measurement limits

Every response of 200 tokens or more counts: shorter ones produce outliers in the log of up to 1,000,000 t/s, one token in almost zero milliseconds. Percentiles by nearest rank, not interpolated.

3

Verifiability

If you do not believe the numbers, clone the repository and run verify_runs.py. It recomputes every value from the raw log and reports any deviation. Currently: none.

Some runs unfortunately do not carry every figure. The Clair Obscure run of Qwen3.6-27B and the three cloud models therefore carry no speed. The artefacts exist, the measurement does not. verify_runs.py logs/