Tests and first-hand experience
Local AI on affordable hardware
Which models deliver which results, and how fast? No standard benchmarks, but image recognition and agentic coding with VS Code tried out first hand on local, affordable AMD hardware.
Tests and first-hand experience
Which models deliver which results, and how fast? No standard benchmarks, but image recognition and agentic coding with VS Code tried out first hand on local, affordable AMD hardware.
Human in the loop
I am Kai Bennett, a software architect, and I have been enthusiastic about local AI since September 2025. Here I share my findings and the results I measured on the two devices.
Personal recommendation
With these two models I ran the longest agentic autonomous jobs, of up to 13.6 h, and at the same time got very good results at a usable median speed of over 20 t/s. For me, on this hardware and as of September 2026, that is the best compromise between quality and speed, without paying for tokens, without depending on cloud providers and without sharing my data.
AI MAX 395
Qwen3.8-Flash-Next
UD-Q4_K_XL
21.76t/s Median
R9700
Qwen3.8-27B
UD-Q4_K_XL
33.69t/s Median
AMD AI MAX 395 · AMD Halo · 13.6 h
Qwen3.8-Flash-Next builds a finished game of — lines of code including tests, autonomously, from a single prompt.
Radeon AI PRO R9700
Qwen3.8-27B UD-Q4_K_XL at over 30 t/s on average, and still around 30 at long context. That is a genuinely usable speed, and this is a dense model in which all 27 billion parameters are active at once.
Machine 1
gfx1201 · RDNA4 · 32 GB dedicated, run headless.
Intel Core Ultra 7 265KF · 63,6 GiB RAM · ASUS PRIME Z890-P.
Windows 11 Pro 26200 · Adrenalin 26.6.4 · amdvlk64 9.2.10.395
Machine 2
Strix Halo · Radeon 8060S · gfx1151 · 128 GiB unified LPDDR5X-8533, around 256 GB/s theoretical bandwidth.
Windows 11 Pro 26200 · Driver 32.0.31021.5001 · Vulkan heap 98.123 MiB
Software
Every run on llama.cpp with the Vulkan backend, no ROCm, no BIOS changes. Served on the local network and wired into VS Code Copilot Chat as a custom endpoint.
Builds b10717 · bd9bd1b · b9985 · 580e88d · poolside 04b2b72
These two models gave me the longest autonomous agent jobs, up to 13.6 h, while still producing very good results at a usable median speed of over 20 t/s. On this hardware, as of September 2026, that is the best compromise I have found.
AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2
UD-Q4_K_XL · MTP, Shared-Q8_0, n-max 2
Not the fastest and not the highest quality either, but it makes use of the advantages of the AMD Strix Halo, meaning it uses the entire VRAM, already includes the Qwen 4 architecture and delivers well at a steady speed.
Radeon AI PRO R9700
UD-Q4_K_XL · draft-mtp + ngram-mod 24/48/64
The first model I would also recommend for coding on a GPU with 24–32 GB of VRAM. In Q4 it runs stably and agentically for hours, writes tests and reviews itself. If I were to recommend one model in the range up to 32 GB of VRAM, it would definitely be this one.
The models
For example a test inspired by the retro game “Moorhuhn”, created from just one prompt, playable for yourself. Every measurement including the speed and the tokens consumed. Image recognition is tested too, an important ability for the models so they can look at their own results and improve them.
For this, an agent was configured in Visual Studio Code carrying the corresponding instructions.
14 of 14 detail pages built·0 to come
Type to filter · ↑ ↓ to select · Enter opens the detail page
Each model appears once, with every task it worked on. Where several runs exist, the median
is computed across all scored responses of those runs together, not averaged from the
individual medians; tokens and minutes are sums. The per-run figures are on the detail
page.
M tokens used is what a paid provider would have charged for: input plus output,
with every single request counted at its full context, not just at what was new. An agent
run resends the same history hundreds of times, which is why input runs a multiple above
output. Short responses are included, even the ones that do not enter the speed figure.
The split into input and output is in each row's tooltip.
The best result so far
Qwen3.8-Flash-Next on the AMD AI MAX 395, 177 billion parameters, roughly six of them active. 459 requests, 786,506 tokens generated, not a single intervention. Graphics, menus, sound, backend: all of it generated along the way. The large models can do this too, of course, but in places I found the result even better than Sonnet 5's, and it runs on hardware costing around 2,600 € (a Bosgame M5, for instance).
Built by Qwen3.8-Flash-Next · AMD AI MAX 395
The result
Recordings from the delivered builds: menus, game modes, scoring, unlocks.
Reference run
Qwen3.8-Flash-Next on the AMD AI MAX 395 (AMD Halo), 177 billion parameters in total, roughly six of them active. Combat arena with magic circle, action menu, status bars and spell effects.
Agentic development
The model runs on the local network and is registered in VS Code Copilot Chat as a custom endpoint, exactly where most developers already are. With tool use it writes tests, creates files and folders, and produces screenshots and Playwright tests to check its own work.
The run is finished: 13 hours 36 minutes on a single prompt, 459 requests, 786,506 tokens generated and 1,051,292 read. Counting cached context, the model read 24.9 million tokens. The context was never truncated once. The peak was 98,109 of 131,072.
The screenshot shows an intermediate state after 8 hours 47 minutes: the seventh of thirteen work packages, 40 files changed, the context window at 51 per cent.
Vision
Many local models can read images. Testing uses my own photographs, to find out how well image recognition works.
In each image exactly one place breaks the pattern. That is a challenge for the models, one where they have to make out detail. Those four places are therefore reported separately.
| What is tested | Qwen3.8-27BUD-Q4_K_XL · Radeon AI PRO R9700 | Qwen3.6-27BUD-Q6_K_XL · Radeon AI PRO R9700 | Qwen3.5-122B-A10BUD-Q4_K_XL · AMD AI MAX 395 (AMD Halo) · GMKtec EVO-X2 | Sonnet 5— · Cloud reference |
|---|---|---|---|---|
| Values read correctly | 44 of 50 | 50 of 50 | 30 of 50 | 46 of 50 |
| Row Terminal-Bench 2.01 | recognised | recognised | recognised | recognised |
| Tile value 40.72 | recognised | recognised | not recognised | not recognised |
| Price tag JOKER3 | not recognised | recognised | not recognised | recognised |
| White mask, unnamed4 | not recognised | recognised | not recognised | order spotted, amounts swapped |
| Price tags fully correct5 | 2 of 4 | 4 of 4 | 0 of 4 | 1 of 4 |
| Invented tags6 | 4 · 1 claimed as certain | 0 | 0 | 5 · none claimed as certain |
Provisional. The answer key for the price tags is confirmed for four of seven tags. Precision and recall are computed against those four only and will change once the rest is confirmed. The four traps are unaffected, they are confirmed.
Speed
33.7t/s
Median over 82 responses. What counts is not the peak speed reached, but the speed you get on average while working through a task.
Context depth
Because they almost always measure at short context. In real use the context grows continuously and speed drops steadily. That is why this page shows median speeds and how they develop with context size.
Prefill · Qwen3.8-27B UD-Q6 · R9700 Report
Decode · Qwen3.8-Flash-Next · AMD AI MAX 395 (AMD Halo) Raw log
On the R9700 the prefill rate drops by a factor of 4.2 between 8 K and 164 K. On the right, the decode rate of Qwen3.8-Flash-Next from the Moorhuhn run: 346 responses of 200 tokens or more, each with its context depth from the log, in six equally sized groups, measured between 34 K and 71 K of context, the range that matters in everyday use. Across that span it drops by a factor of 1.18. The model ran especially steadily and reliably.
Fine-tuning
The parameters in the model cards are not always optimal (for AMD). These three, for example, I changed.
spec-draft-n-max · Laguna S 2.1
15 → 3
With the documented --spec-draft-n-max 15 decode falls to
8.1 t/s, below the 20.55 achieved without any speculation at all. With
3 it rises to 27.9.
Sampling · Qwen3.5-122B
0.0 → 1.5
With presence_penalty 0.0 one task ran to the end of the context:
32,768 tokens, never finishing. The same task needs 201 tokens with the
setting switched on.
Memory · Strix Halo
25.6 GiB
A larger UMA reservation brought no gain and would have taken memory away from the operating system. Even at 262,144 context, 25.6 GiB remained free.
Model cards
Every row is a script that started exactly one such server. Each set of parameters can be copied on its own, complete, with every flag.
| Model | Hardware | Context | -ub | KV k / v | Speculation | Vision | |
|---|---|---|---|---|---|---|---|
| Qwen3.6-27B-UD-Q6_K_XLstart_qwen36_27b_q6_coding.bat | R9700 | 163,840 | 512 | q8_0 / q4_0 | — | Vision | |
| Qwen3.6-27B-UD-Q6_K_XL-mtpstart_qwen36_27b_q6_mtp_coding.bat | R9700 | 172,032 | 512 | q8_0 / q4_0 | draft-mtp | Vision | |
| start_qwen36_35b_a3b_mtp.bat | R9700 | 172,032 | — | — / — | — | — | |
| start_qwen38_27b_coding.bat | R9700 | 262,144 | — | — / — | — | — | |
| start_qwen38_27b_ngram_test.bat | R9700 | 262,144 | — | — / — | — | — | |
| start_qwen38_27b_q4xl_coding.bat | R9700 | 262,144 | — | — / — | — | — | |
| start_qwen38_27b_q4xl_medium.bat | R9700 | 262,144 | — | — / — | — | — | |
| Qwen3.8-27B-UD-Q6_K_Mstart_qwen38_27b_test_rmkq1_ub128.bat | R9700 | 262,144 | 128 | q8_0 / turbo4 | draft-mtp | Vision | |
| Qwen3.8-27B-UD-Q6_K_Mstart_qwen38_27b_test_ub256.bat | R9700 | 262,144 | 256 | q8_0 / turbo4 | draft-mtp | Vision | |
| Qwen3.8-27B-UD-Q6_K_Mstart_qwen38_27b_test_ub64.bat | R9700 | 262,144 | 64 | q8_0 / turbo4 | draft-mtp | Vision | |
| DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004start-deepseek-v4-flash-0731-DSPARK.bat | AI MAX 395 | 131,072 | 512 | f16 / f16 | draft-dspark | — | |
| Qwen3.6-35B-A3B-UD-Q4_K_XLstart-qwen-3.6-35b-a3b-vision.bat | AI MAX 395 | 262,144 | — | — / — | — | Vision | |
| Qwen3.5-122B-A10B-UD-Q4_K_XL-00001-of-00003start-qwen35-122b-A10B-MTP-NOTHINK.bat | AI MAX 395 | 262,144 | 512 | q8_0 / q8_0 | draft-mtp | Vision | |
| Qwen3.5-122B-A10B-UD-Q4_K_XL-00001-of-00003start-qwen35-122b-A10B-MTP.bat | AI MAX 395 | 262,144 | 512 | q8_0 / q8_0 | draft-mtp | Vision | |
| Qwen3.6-27B-UD-Q6_K_XLstart-qwen36-27b-MTP.bat | AI MAX 395 | 262,144 | 512 | q8_0 / q4_0 | draft-mtp | Vision | |
| Qwen3.8-27B-UD-Q6_K_XLstart-qwen38-27b-FALLBACK-8090.bat | AI MAX 395 | 32,768 | 512 | q8_0 / q4_0 | draft-mtp | Vision | |
| Qwen3.8-27B-UD-Q6_K_XL-240819start-qwen38-27b-MTP-240819.bat | AI MAX 395 | 262,144 | 512 | q8_0 / q4_0 | draft-mtp | Vision | |
| Qwen3.8-27B-UD-Q6_K_XLstart-qwen38-27b-MTP.bat | AI MAX 395 | 262,144 | 512 | q8_0 / q4_0 | draft-mtp | Vision | |
| Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004start-qwen38-flash-next.bat | AI MAX 395 | 131,072 | 512 | q8_0 / q8_0 | draft-mtp | Vision |
Agent Test Harness
Two different things are measured. The run shows speed, stamina and whether something runnable exists at the end. The detail page shows: design ability, image recognition, and whether it holds to a contract without inventing.
01
Two prompts, from which every result on this page came. Unchanged, exactly as the models received them. They are deliberately built as opposites.
One prompt can be passed by diligence, the other cannot.
02
One agent for VS Code
Copilot Chat. Drop it into
.github/agents/ and it shows
up in the chat picker.
llama-server speaks the OpenAI API. VS Code accepts such endpoints through “bring your own key”, without a Copilot subscription and without signing in. The launch parameters for each run are given under model cards.
03
Every run has its own detail page with the same eleven sections: architecture, measurements, artefact, image recognition, speed, configuration, quality, errors, sources, reproduction.
Test bench
Copy the launch parameters of the fastest measured run on each machine.
gfx1201 · RDNA4 · 32 GB dedicated
from ~1,500 €
Fastest run: Qwen3.8-27B UD-Q4_K_XL · 33.69 t/s
Strix Halo · Radeon 8060S · gfx1151 · 128 GiB unified
from ~1,800 €
Fastest run: Qwen3.8-Flash-Next UD-Q4_K_XL · 21.76 t/s
Three numbers for the same memory. The AMD AI MAX 395 has 128 GiB of unified LPDDR5X, not 128 GB of VRAM. Windows sees 63.6 GiB, the BIOS reservation is 64.4 GiB, and Vulkan reports 98,123 MiB.
Verifiability
Logs and measurements are in the public repository and were taken on the two systems described above. Depending on your exact hardware and software setup, your numbers may of course differ somewhat.
1
The llama.cpp server logs and the measurement tables sit unchanged in
evidence/, around 3.8 MB in total, with checksums. Every model row
and every table line links exactly the file its numbers come from.
2
Every response of 200 tokens or more counts: shorter ones produce outliers in the log of up to 1,000,000 t/s, one token in almost zero milliseconds. Percentiles by nearest rank, not interpolated.
3
If you do not believe the numbers, clone the repository and run
verify_runs.py. It recomputes every value from the raw log and
reports any deviation. Currently: none.
Some runs unfortunately do not carry every figure. The Clair Obscure run of Qwen3.6-27B and the three cloud models therefore carry no speed. The artefacts exist, the measurement does not. verify_runs.py logs/
Source code
First the repository behind this site: it holds every log, every piece of evidence and the source code of the benchmarks. Below it the tools that were in use during these measurements.
The repository behind this site. evidence/ holds the server logs,
chat transcripts, image-test answers and measurement reports, with checksums;
benchmarks/ holds the source code of every game that was built,
sorted by model and run. Every evidence link on this page points there.
Hermes Agent and Claude Code running entirely locally through llama.cpp. A session of 4 hours and 7 million tokens would have cost around 94 dollars in the cloud.
Gemma-4-31B with a full 256 K context on an RDNA4 card, TurboQuant KV cache and flash attention for llama.cpp, with real measurements.
Fine-tuning on Radeon GPUs with QLoRA over ROCm, on Windows via WSL2 and on Linux. With a Gemma-4 example, a live dashboard and validation.
The llama.cpp fork with the TurboQuant KV cache on which part of these
measurements was produced, build bd9bd1b.