Cloud reference
Opus 5 ULTRACODE
Cloud reference
- Hardware
- Cloud reference—
- Memory
- —
- Tasks
- Clair Obscure, Image recognition
- Quantisations
- —
About the model
Context window of 1 million tokens, which is both the default and the maximum; the documentation states there is no smaller context variant. Input is text and images, output is text, up to 128,000 tokens synchronously and up to 300,000 tokens via the Message Batches API with the output-300k-2026-03-24 beta header. Adaptive thinking is on by default and is steered by the effort parameter across five levels from low to max, the default is high, and thinking cannot be disabled at xhigh or max. Anthropic's model documentation states neither a parameter count nor a layer count or architecture type; what is documented is the model ID claude-opus-5, the release date of July 24, 2026, and a training data cutoff and reliable knowledge cutoff of May 2026.
- Model ID
- claude-opus-5
- Context window
- 1M tokens (default and maximum)
- Maximum output (Messages API)
- 128K tokens
- Maximum output (Message Batches API, beta)
- 300K tokens
- Modalities
- Text and images in, text out
- Reasoning
- Adaptive thinking, on by default
- Effort levels
- low, medium, high, xhigh, max
- Default effort
- high
Figures from the vendor: Source · Model card · Vendor
The numbers at a glance
The run went through a third-party interface. There is no server log, so neither speed nor token balance. What counts is the artefact.
Artefact available and playable. Speed in the cloud cannot be measured comparably.
Results
No playable build is on record for this run.
Assessment
34 / 40
The grade is a personal, and therefore subjective, assessment.
Game feelDoes it run, and does it play?
10/10
Playable, the most complete implementation of the brief in the field.
PresentationMenus, HUD, graphics, sound
10/10
Recording with loading screen, combat and menus.
Code qualityTests, structure, self-corrections
4/10
No test files.
ScopeHow much of the prompt was fulfilled?
10/10
58 files, 16,308 lines, by far the largest scope.
This assessment is a proposal and has not yet been confirmed by Kai Bennett. Code quality and scope rest on counted files, source and test lines and repair scripts in the evidence repository.
How do you see it? The grade above is a personal assessment. Here yours counts, independently of it.
What went wrong
- The game loads three GLB models and a 4K HDRI on demand; on slow connections the first start takes over ten seconds.
- One resource is missing and returns 404 — visible in the browser console, with no effect on the picture.
- The controls require a keyboard. On touchscreens the game cannot be played.
Artefact available and playable. Speed in the cloud cannot be measured comparably.
Image recognition
The answer key for the vision test was written by Opus 5 itself. A hit rate of its own would be circular and is therefore not reported. The redraw of the chart is shown here anyway — as an example of what is expected in this section.
Sources
- Model carddocs.anthropic.com/en/docs/about-claude/models
- Vendor sitewww.anthropic.com
- Measurement repositorygithub.com/KaiFelixBennett/local-ai-amd-benchmark
- Source code · Clair Obscuregithub.com/KaiFelixBennett/local-ai-amd-benchmark/tree/main/benchmarks/opus5/clairobscure
Try it yourself
- Download the weights
- Start llama-server with the parameters above
- Register the endpoint in VS Code as a custom model and pick the agent
Prompt, agent files and tools: Agent Test Harness
Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run opus5-clairobscure. Measurement data CC-BY-4.0.
Measurement data CC-BY-4.0, code MIT.