Skip to content

Cloud reference

Sonnet 5

Cloud reference

Hardware
Cloud reference—
Memory
—
Tasks
Moorhuhn, Image recognition
Quantisations
—

About the model

Anthropic publishes no architecture data for Claude Sonnet 5. Neither the Claude platform model page nor the system card of June 30, 2026 states a parameter count, a layer count, or whether the model is dense or built as a mixture of experts. Only operating data is documented: a 1 million token context window, 128,000 tokens maximum output, text and images as input, text as output, adaptive thinking with five effort levels, and a training data cutoff in January 2026. The system card describes the training as a proprietary mix of publicly available information from the internet, public and private datasets, and synthetic data generated by other models, with no figures for compute or data volume.

Model ID
claude-sonnet-5
Context window
1,000,000 tokens
Maximum output
128,000 tokens
Maximum output (Batch API, beta)
300,000 tokens
Input and output
Text and images to text
Thinking mode
Adaptive thinking, on by default
Effort levels
low, medium, high, xhigh, max; default high
Tokenizer
About 30 percent more tokens per text than Sonnet 4.6

Figures from the vendor: Source · Model card · Vendor

The numbers at a glance

The run went through Claude Code. There is no server log, only the session log. The working time comes from it; a speed cannot be stated. What counts is the artefact.

Runtime

0.8 h

42 minutes of work and 7 minutes waiting at the subscription's session limit

Self-corrections

2

Cloud reference in Claude Code with effort high. The artefact is available; speed is not measured comparably. The working time of 42 minutes is calculated from the session log. Added to that are 7 minutes of waiting at the session limit of the Claude subscription (reached at 11:32, cleared from 11:40), so 49 minutes until the finished game. The 5 minutes after that until the author typed “continue” are deducted.

Results

Moorland Mayhem: recording of the build Sonnet 5 delivered.

Moorland Mayhem — Featherstorm

Opens in its own layer

Aim and shoot with mouse or finger. Needs WebGL.

Loads only on click · 0.3 MB transfer · 1.6 MB unpacked

Assessment

30 / 40

The grade is a personal, and therefore subjective, assessment.

Game feelDoes it run, and does it play?

Playable, but the sound is poor, for example the droning.

PresentationMenus, HUD, graphics, sound

Looks very childish, otherwise good.

Code qualityTests, structure, self-corrections

Type check, build and ESLint clean, all 100 tests green, cleanly layered. With 885 test lines, though, the thinnest test coverage of the rated runs.

ScopeHow much of the prompt was fulfilled?

Six game modes, three maps and ten screens with bosses, chain reactions and a daily challenge. There is no shop.

What went wrong

There were 2 self-corrections by the model.

Cloud reference in Claude Code with effort high. The artefact is available; speed is not measured comparably. The working time of 42 minutes is calculated from the session log. Added to that are 7 minutes of waiting at the session limit of the Claude subscription (reached at 11:32, cleared from 11:40), so 49 minutes until the finished game. The 5 minutes after that until the author typed “continue” are deducted.

Image recognition

A separate test run, not part of a measured run. Cloud reference.

Values read

46 of 50

dashboard, 50 checkable entries

Price tags

1 of 4

read fully correctly

Invented tags

2

none claimed as certain; 3 reported as unreadable

Stumbling points spotted

2 of 4

places that break the pattern

Provisional: the answer key for the price tags is not fully confirmed yet.

Task 1 · reading values

The test image, a dashboard of my own
The test image. A German-language screenshot of securesight.ai
10203040506070'23'24'25'26NOWIntelligence Index (AA v4.0)GPT-4GPT-4oClaude 3.5 SonnetOpus 4.6GPT-5.5Claude Fable 5Qwen-Next 80BQwen3.6 27B ★Gemma 3 27BGemma 4 12BGemma 4 31B ★Ministral 3 14BMistral Small 4Proprietäre FlaggschiffeQwen openGemma openMistral
What the model redrew from it as SVG.

What the model read from the tooltip table

Claude Opus 4.5Qwen3.6 27B
Intelligence Index40.837.1
GPQA Diamond8784
Humanity's Last Exam2922
Terminal-Bench 2.05061

And from the KPI tiles, in German as on the test image

  • QWEN3.6 27B · OPEN 37.1 · AA Intelligence Index — Open-Weight, kein API-Key, läuft auf ihrer eigenen GPU.
  • VS. OPUS 4.5 (40,8) 91% · des Opus-4.5-Niveaus im Gesamtindex. Bei Coding-Benchmarks: gleichauf.
  • PROPRIETÄRE SPITZE · HEUTE 59.9 · Das Rechenzentrum klettert weiter — doch der Abstand schrumpft jedes Quartal.

Task 2 · complex text recognition

What is tested is the reading of tiny, tilted, partly hidden text in a photo with more than forty objects, plus two currency symbols that are hard to tell apart once rotated. The price tags are what the test is run on.

The second test image, a trade fair stand
The second test image. The model gave no coordinates for its tags, so nothing is marked in it.

What the model read

What the model read there £€ Ordercertain
1 MASS EFFECT not readable not readable · no
2 JOKER 30 40 gbp zuerst no
3 unnamed 45 35 eur zuerst no
4 unnamed not readable not readable · no
5 unnamed not readable not readable · no
6 unnamed not readable not readable · no

The model counted 8 tags in the scored area. The answer key is deliberately not shown beside it, otherwise a later model could read it off this page.

The model on the hardest part, quoted in the German it answered in: „Im Diagramm die eng gedrängten Punktbeschriftungen oben rechts (Qwen-, Gemma- und Mistral-Punkte liegen dicht beieinander, Labels überschneiden sich fast). In Bild B die Preisziffern selbst: die Schilder liegen schräg zur Kamera und sind bei mehreren Masken so klein/unscharf abgebildet, dass ich nicht einmal den vollständigen Produktnamen sicher lesen konnte, geschweige denn die Beträge.“

Where the model called itself unsure (10 places)
  • Bild A: erster X-Achsen-Wert — gelesen als '23, könnte auch '22 sein
  • Bild A: Klammerwert in Kachel 2 — gelesen als '(40,8)', Ziffer nicht eindeutig
  • Bild A: Y-Achsenbeschriftung '(AA v4.0)' im Klammerteil unsicher
  • Bild A: Fußzeile unter der Tooltip-Tabelle nur teilweise lesbar
  • Bild A: mögliche kleine Zahlen-Abzeichen neben KIMI und MISTRAL im Filter — Ziffern zu klein für eine sichere Lesung
  • Bild A: Zusatztext neben dem Mistral-Legendeneintrag nicht sicher entzifferbar
  • Bild B: bei den meisten Schildern im Wertungsbereich waren die Pfund-/Euro-Ziffern zu klein bzw. zu schräg fotografiert, um sie ohne Raten zu übernehmen — deshalb dort null statt einer erfundenen Zahl
  • Bild B: Reihenfolge £/€ bei mehreren Schildern nicht sicher bestimmbar
  • Bild B: genaue Zusatzbezeichnungen der drei 'REDHOOD'-Masken (z. B. V1 / Armor / Battle Damaged / Airman) nicht sicher unterscheidbar
  • Bild B: die Abgrenzung des Wertungsbereichs (x bis 700, y ab 950) konnte ich nur optisch schätzen, ohne Koordinatenwerkzeug — einzelne Schilder nahe der Grenze könnten falsch zugeordnet sein

The model wrote this list itself, not the scorer, and in German; it is quoted as written. It counts: an admitted “not readable” costs one point, an invented entry costs more.

The complete answer of the model as JSON · The scoring is not published beside it: it names the correct value for every mistake, which would put the answer key in the open.

Try it yourself

  1. Download the weights
  2. Start llama-server with the parameters above
  3. Register the endpoint in VS Code as a custom model and pick the agent

Prompt, agent files and tools: Agent Test Harness

Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run sonnet5. Measurement data CC-BY-4.0.

Measurement data CC-BY-4.0, code MIT.