Cloud reference
Sonnet 5
Cloud reference
- Hardware
- Cloud reference—
- Memory
- —
- Tasks
- Moorhuhn, Image recognition
- Quantisations
- —
About the model
Anthropic publishes no architecture data for Claude Sonnet 5. Neither the Claude platform model page nor the system card of June 30, 2026 states a parameter count, a layer count, or whether the model is dense or built as a mixture of experts. Only operating data is documented: a 1 million token context window, 128,000 tokens maximum output, text and images as input, text as output, adaptive thinking with five effort levels, and a training data cutoff in January 2026. The system card describes the training as a proprietary mix of publicly available information from the internet, public and private datasets, and synthetic data generated by other models, with no figures for compute or data volume.
- Model ID
- claude-sonnet-5
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Maximum output (Batch API, beta)
- 300,000 tokens
- Input and output
- Text and images to text
- Thinking mode
- Adaptive thinking, on by default
- Effort levels
- low, medium, high, xhigh, max; default high
- Tokenizer
- About 30 percent more tokens per text than Sonnet 4.6
Figures from the vendor: Source · Model card · Vendor
The numbers at a glance
The run went through Claude Code. There is no server log, only the session log. The working time comes from it; a speed cannot be stated. What counts is the artefact.
Runtime
0.8 h
42 minutes of work and 7 minutes waiting at the subscription's session limit
Self-corrections
2
Cloud reference in Claude Code with effort high. The artefact is available; speed is not measured comparably. The working time of 42 minutes is calculated from the session log. Added to that are 7 minutes of waiting at the session limit of the Claude subscription (reached at 11:32, cleared from 11:40), so 49 minutes until the finished game. The 5 minutes after that until the author typed “continue” are deducted.
Results
Moorland Mayhem — Featherstorm
Opens in its own layerAim and shoot with mouse or finger. Needs WebGL.
Loads only on click · 0.3 MB transfer · 1.6 MB unpacked
Assessment
30 / 40
The grade is a personal, and therefore subjective, assessment.
Game feelDoes it run, and does it play?
7/10
Playable, but the sound is poor, for example the droning.
PresentationMenus, HUD, graphics, sound
8/10
Looks very childish, otherwise good.
Code qualityTests, structure, self-corrections
7/10
Type check, build and ESLint clean, all 100 tests green, cleanly layered. With 885 test lines, though, the thinnest test coverage of the rated runs.
ScopeHow much of the prompt was fulfilled?
8/10
Six game modes, three maps and ten screens with bosses, chain reactions and a daily challenge. There is no shop.
How do you see it? The grade above is a personal assessment. Here yours counts, independently of it.
What went wrong
There were 2 self-corrections by the model.
Cloud reference in Claude Code with effort high. The artefact is available; speed is not measured comparably. The working time of 42 minutes is calculated from the session log. Added to that are 7 minutes of waiting at the session limit of the Claude subscription (reached at 11:32, cleared from 11:40), so 49 minutes until the finished game. The 5 minutes after that until the author typed “continue” are deducted.
Image recognition
A separate test run, not part of a measured run. Cloud reference.
Values read
46 of 50
dashboard, 50 checkable entries
Price tags
1 of 4
read fully correctly
Invented tags
2
none claimed as certain; 3 reported as unreadable
Stumbling points spotted
2 of 4
places that break the pattern
Provisional: the answer key for the price tags is not fully confirmed yet.
Task 1 · reading values
What the model read from the tooltip table
| Claude Opus 4.5 | Qwen3.6 27B | |
|---|---|---|
| Intelligence Index | 40.8 | 37.1 |
| GPQA Diamond | 87 | 84 |
| Humanity's Last Exam | 29 | 22 |
| Terminal-Bench 2.0 | 50 | 61 |
And from the KPI tiles, in German as on the test image
- QWEN3.6 27B · OPEN 37.1 · AA Intelligence Index — Open-Weight, kein API-Key, läuft auf ihrer eigenen GPU.
- VS. OPUS 4.5 (40,8) 91% · des Opus-4.5-Niveaus im Gesamtindex. Bei Coding-Benchmarks: gleichauf.
- PROPRIETÄRE SPITZE · HEUTE 59.9 · Das Rechenzentrum klettert weiter — doch der Abstand schrumpft jedes Quartal.
Task 2 · complex text recognition
What is tested is the reading of tiny, tilted, partly hidden text in a photo with more than forty objects, plus two currency symbols that are hard to tell apart once rotated. The price tags are what the test is run on.
What the model read
| What the model read there | £ | € | Order | certain | |
|---|---|---|---|---|---|
| 1 | MASS EFFECT | not readable | not readable | · | no |
| 2 | JOKER | 30 | 40 | gbp zuerst | no |
| 3 | unnamed | 45 | 35 | eur zuerst | no |
| 4 | unnamed | not readable | not readable | · | no |
| 5 | unnamed | not readable | not readable | · | no |
| 6 | unnamed | not readable | not readable | · | no |
The model counted 8 tags in the scored area. The answer key is deliberately not shown beside it, otherwise a later model could read it off this page.
The model on the hardest part, quoted in the German it answered in: „Im Diagramm die eng gedrängten Punktbeschriftungen oben rechts (Qwen-, Gemma- und Mistral-Punkte liegen dicht beieinander, Labels überschneiden sich fast). In Bild B die Preisziffern selbst: die Schilder liegen schräg zur Kamera und sind bei mehreren Masken so klein/unscharf abgebildet, dass ich nicht einmal den vollständigen Produktnamen sicher lesen konnte, geschweige denn die Beträge.“
Where the model called itself unsure (10 places)
- Bild A: erster X-Achsen-Wert — gelesen als '23, könnte auch '22 sein
- Bild A: Klammerwert in Kachel 2 — gelesen als '(40,8)', Ziffer nicht eindeutig
- Bild A: Y-Achsenbeschriftung '(AA v4.0)' im Klammerteil unsicher
- Bild A: Fußzeile unter der Tooltip-Tabelle nur teilweise lesbar
- Bild A: mögliche kleine Zahlen-Abzeichen neben KIMI und MISTRAL im Filter — Ziffern zu klein für eine sichere Lesung
- Bild A: Zusatztext neben dem Mistral-Legendeneintrag nicht sicher entzifferbar
- Bild B: bei den meisten Schildern im Wertungsbereich waren die Pfund-/Euro-Ziffern zu klein bzw. zu schräg fotografiert, um sie ohne Raten zu übernehmen — deshalb dort null statt einer erfundenen Zahl
- Bild B: Reihenfolge £/€ bei mehreren Schildern nicht sicher bestimmbar
- Bild B: genaue Zusatzbezeichnungen der drei 'REDHOOD'-Masken (z. B. V1 / Armor / Battle Damaged / Airman) nicht sicher unterscheidbar
- Bild B: die Abgrenzung des Wertungsbereichs (x bis 700, y ab 950) konnte ich nur optisch schätzen, ohne Koordinatenwerkzeug — einzelne Schilder nahe der Grenze könnten falsch zugeordnet sein
The model wrote this list itself, not the scorer, and in German; it is quoted as written. It counts: an admitted “not readable” costs one point, an invented entry costs more.
The complete answer of the model as JSON · The scoring is not published beside it: it names the correct value for every mistake, which would put the answer key in the open.
Sources
- Model carddocs.anthropic.com/en/docs/about-claude/models
- Vendor sitewww.anthropic.com
- Measurement repositorygithub.com/KaiFelixBennett/local-ai-amd-benchmark
- Source code · Moorhuhngithub.com/KaiFelixBennett/local-ai-amd-benchmark/tree/main/benchmarks/sonnet5/moorhuhn
Try it yourself
- Download the weights
- Start llama-server with the parameters above
- Register the endpoint in VS Code as a custom model and pick the agent
Prompt, agent files and tools: Agent Test Harness
Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run sonnet5. Measurement data CC-BY-4.0.
Measurement data CC-BY-4.0, code MIT.