Skip to content

Cloud reference

Opus 5

Cloud reference

Hardware
Cloud reference—
Memory
—
Tasks
Moorhuhn
Quantisations
—

About the model

Context window of 1 million tokens, which is both the default and the maximum; the documentation states there is no smaller context variant. Input is text and images, output is text, up to 128,000 tokens synchronously and up to 300,000 tokens via the Message Batches API with the output-300k-2026-03-24 beta header. Adaptive thinking is on by default and is steered by the effort parameter across five levels from low to max, the default is high, and thinking cannot be disabled at xhigh or max. Anthropic's model documentation states neither a parameter count nor a layer count or architecture type; what is documented is the model ID claude-opus-5, the release date of July 24, 2026, and a training data cutoff and reliable knowledge cutoff of May 2026.

Model ID
claude-opus-5
Context window
1M tokens (default and maximum)
Maximum output (Messages API)
128K tokens
Maximum output (Message Batches API, beta)
300K tokens
Modalities
Text and images in, text out
Reasoning
Adaptive thinking, on by default
Effort levels
low, medium, high, xhigh, max
Default effort
high

Figures from the vendor: Source · Model card · Vendor

The numbers at a glance

The run went through Claude Code. There is no server log, only the session log. Working time and token count come from it; a speed cannot be stated.

Runtime

3.6 h

112 minutes of work and 102 minutes waiting at the subscription's session limit

Cloud reference in Claude Code with effort high, not in ULTRACODE mode: the session log shows effort high in all 230 requests and no Ultracode activation. Working time and tokens are calculated from the log.

Assessment

not yet assessed

The grade is a personal, and therefore subjective, assessment.

There is no artefact to assess for this run, so it carries no grade.

What went wrong

No self-corrections were logged.

Cloud reference in Claude Code with effort high, not in ULTRACODE mode: the session log shows effort high in all 230 requests and no Ultracode activation. Working time and tokens are calculated from the log.

Try it yourself

  1. Download the weights
  2. Start llama-server with the parameters above
  3. Register the endpoint in VS Code as a custom model and pick the agent

Prompt, agent files and tools: Agent Test Harness

Citation: Kai Felix Bennett, “Usability test of local AI on AMD hardware”, benchmark.securesight.ai, run opus5-moorhuhn. Measurement data CC-BY-4.0.

Measurement data CC-BY-4.0, code MIT.