Skip to content
benchmark.securesight.ai
DE

Cloud reference

Opus 5 ULTRACODE

Cloud reference · 2 runs

Hardware
Cloud reference
Memory
Tasks
Clair Obscure, Moorhuhn, Image recognition
Quantisations

About the model

Context window of 1 million tokens, which is both the default and the maximum; the documentation states there is no smaller context variant. Input is text and images, output is text, up to 128,000 tokens synchronously and up to 300,000 tokens via the Message Batches API with the output-300k-2026-03-24 beta header. Adaptive thinking is on by default and is steered by the effort parameter across five levels from low to max, the default is high, and thinking cannot be disabled at xhigh or max. Anthropic's model documentation states neither a parameter count nor a layer count or architecture type; what is documented is the model ID claude-opus-5, the release date of July 24, 2026, and a training data cutoff and reliable knowledge cutoff of May 2026.

Modell-ID
claude-opus-5
Kontextfenster
1 Mio. Token (Standard und Maximum)
Maximale Ausgabe (Messages API)
128K Token
Maximale Ausgabe (Message Batches API, Beta)
300K Token
Modalitäten
Text und Bilder rein, Text raus
Reasoning
Adaptives Denken, standardmäßig aktiv
Effort-Stufen
low, medium, high, xhigh, max
Standard-Effort
high

Figures from the vendor: Source · Model card · Vendor

Run: Clair Obscure · 25 July 2026

The numbers at a glance

Der Lauf ging über eine fremde Schnittstelle. Es gibt kein Serverprotokoll, also weder Geschwindigkeit noch Tokenbilanz. Was zählt, ist das Artefakt.

Artefakt vorhanden und spielbar. Geschwindigkeit in der Cloud nicht vergleichbar messbar.

Results

Gilded Requiem — Ladebildschirm, Kampfarena und Dialogszene, ungeschnitten
Der Ladebildschirm: drei Figuren mit Werten, Schwierigkeitsgrad, elf Luminas zur Auswahl
Der Ladebildschirm: drei Figuren mit Werten, Schwierigkeitsgrad, elf Luminas zur Auswahl
Le Parvis Noyé — die Arena mit schwebenden Trümmern, Kronleuchtern und Dialogzeile
Le Parvis Noyé — die Arena mit schwebenden Trümmern, Kronleuchtern und Dialogzeile
Rundenkampf mit Aktionsmenü und Zeitfenster für Parade und Ausweichen
Rundenkampf mit Aktionsmenü und Zeitfenster für Parade und Ausweichen

No playable build is on record for this run.

Assessment

34 / 40

The grade is a personal, and therefore subjective, assessment.

Game feelDoes it run, and does it play?

Spielbar, die vollständigste Umsetzung des Auftrags im Feld.

PresentationMenus, HUD, graphics, sound

Aufnahme mit Ladebildschirm, Kampf und Menüs.

Code qualityTests, structure, self-corrections

Keine Testdateien.

ScopeHow much of the prompt was fulfilled?

58 Dateien, 16 308 Zeilen, mit Abstand der größte Umfang.

This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.

What went wrong

  • Das Spiel lädt drei GLB-Modelle und eine 4K-HDRI nach; auf schwachen Verbindungen dauert der erste Start über zehn Sekunden.
  • Eine Ressource fehlt und liefert 404 — sichtbar in der Browserkonsole, ohne Wirkung auf das Bild.
  • Die Steuerung setzt Tastatur voraus. Auf Tastbildschirmen ist das Spiel nicht bedienbar.

Artefakt vorhanden und spielbar. Geschwindigkeit in der Cloud nicht vergleichbar messbar.

Run: Moorhuhn · 02 September 2026

The numbers at a glance

Dieser Lauf steht noch aus. Es liegt nur der Auftrag vor, keine Messung und kein Artefakt.

Kein Lauf, keine Zahlen.

Assessment

not yet assessed

The grade is a personal, and therefore subjective, assessment.

There is no artefact to assess for this run, so it carries no grade.

What went wrong

No self-corrections were logged.

Kein Lauf, keine Zahlen.

Image recognition

Der Lösungsschlüssel für den Bilderkennungs-Test wurde von Opus 5 selbst geschrieben. Eine eigene Trefferquote wäre zirkulär und wird deshalb nicht ausgewiesen. Der Nachbau des Diagramms steht trotzdem hier — als Beispiel dafür, was in diesem Abschnitt erwartet wird.

Try it yourself

  1. Download the weights
  2. Start llama-server with the parameters above
  3. Register the endpoint in VS Code as a custom model and pick the agent

Prompt, agent files and tools: Agent Test Harness

Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run opus5-clairobscure. Measurement data CC-BY-4.0.

Measurement data CC-BY-4.0, code MIT.