Cloud reference
Opus 5
Cloud reference
- Hardware
- Cloud reference—
- Memory
- —
- Tasks
- Moorhuhn
- Quantisations
- —
About the model
Context window of 1 million tokens, which is both the default and the maximum; the documentation states there is no smaller context variant. Input is text and images, output is text, up to 128,000 tokens synchronously and up to 300,000 tokens via the Message Batches API with the output-300k-2026-03-24 beta header. Adaptive thinking is on by default and is steered by the effort parameter across five levels from low to max, the default is high, and thinking cannot be disabled at xhigh or max. Anthropic's model documentation states neither a parameter count nor a layer count or architecture type; what is documented is the model ID claude-opus-5, the release date of July 24, 2026, and a training data cutoff and reliable knowledge cutoff of May 2026.
- Model ID
- claude-opus-5
- Context window
- 1M tokens (default and maximum)
- Maximum output (Messages API)
- 128K tokens
- Maximum output (Message Batches API, beta)
- 300K tokens
- Modalities
- Text and images in, text out
- Reasoning
- Adaptive thinking, on by default
- Effort levels
- low, medium, high, xhigh, max
- Default effort
- high
Figures from the vendor: Source · Model card · Vendor
The numbers at a glance
The run went through Claude Code. There is no server log, only the session log. Working time and token count come from it; a speed cannot be stated.
Runtime
3.6 h
112 minutes of work and 102 minutes waiting at the subscription's session limit
Cloud reference in Claude Code with effort high, not in ULTRACODE mode: the session log shows effort high in all 230 requests and no Ultracode activation. Working time and tokens are calculated from the log.
Assessment
not yet assessed
The grade is a personal, and therefore subjective, assessment.
There is no artefact to assess for this run, so it carries no grade.
What went wrong
No self-corrections were logged.
Cloud reference in Claude Code with effort high, not in ULTRACODE mode: the session log shows effort high in all 230 requests and no Ultracode activation. Working time and tokens are calculated from the log.
Sources
- Model carddocs.anthropic.com/en/docs/about-claude/models
- Vendor sitewww.anthropic.com
- Measurement repositorygithub.com/KaiFelixBennett/local-ai-amd-benchmark
Try it yourself
- Download the weights
- Start llama-server with the parameters above
- Register the endpoint in VS Code as a custom model and pick the agent
Prompt, agent files and tools: Agent Test Harness
Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run opus5-moorhuhn. Measurement data CC-BY-4.0.
Measurement data CC-BY-4.0, code MIT.