Anthropic released Claude Opus 5 on July 24, 2026, and the detail most relevant to research institutions is not a benchmark score but a configuration change: reasoning effort is now a per-query dial rather than a per-model choice. Where earlier generations of frontier models were selected once, at procurement time, for a fixed cost/capability tradeoff, Claude Opus 5 exposes an effort parameter with multiple settings — low, medium, high (the default), xhigh, and, in some evaluation contexts, an additional max setting — that a developer or institutional integrator sets per call. The same underlying model can run cheap and fast for routine work or slow and expensive for the hardest task in the same afternoon, without switching models or renegotiating a license.
What changed, in Anthropic’s own terms
Anthropic’s developer documentation describes the tradeoff plainly: low and medium effort “produce strong quality at a fraction of the tokens and latency of higher settings,” and the recommended pattern is to use the lower tiers as the primary lever for token cost and response time wherever quality holds, stepping up to xhigh only for demanding coding and agentic work. Thinking is enabled by default and can only be disabled at high effort or below — a departure from Claude Opus 4.8, where thinking was optional at every tier. Pricing itself did not change with the tier system: Opus 5 lists at the same $5-per-million-input-token / $25-per-million-output-token rate as Opus 4.8, so the cost variable institutions now manage is effort-driven token and latency consumption within a single flat per-token price, not a menu of different model prices. A separate fast mode runs at roughly 2.5x the default speed for latency-sensitive use.
Independent benchmarking gives a sense of where the ceiling sits. On Artificial Analysis’s Intelligence Index, Claude Opus 5 at Max and xHigh effort tied at the top of the board (63), with High effort close behind (61); on the same firm’s Agentic Index, Opus 5 at Max effort tied for first with Grok 4.6 at High effort (59). The practical read for a research-computing or AI-governance office is that the top tiers buy a real, measurable capability increase over the default — but the default tier is already close enough on most tasks that Anthropic’s own guidance is to reserve the expensive settings for cases that actually need them.
Why this matters for institutional AI procurement
Most institutional AI-tool policies to date have been written around model-level decisions: which vendor, which model, what data-handling terms, what per-seat or per-token budget. An effort-tiered model complicates that framing in a specific way — the same contractual relationship and the same model can now produce meaningfully different cost, latency, and output-quality profiles depending on a setting an end user, a vendor’s default configuration, or an integrating tool chooses on their behalf, often invisibly. A researcher using a university-licensed AI writing or coding assistant built on Claude Opus 5 may have no visibility into, or control over, which effort tier is actually running behind the interface. For a research-computing or AI-governance office evaluating and licensing LLM tools, this argues for asking vendors a more granular question than “which model do you use”: which effort tier is the default in this integration, is it configurable, and does the vendor disclose when it silently downgrades effort under load or cost pressure.
The agentic dimension raises the governance stakes further. Anthropic’s own prompting guidance for Opus 5 notes that the model delegates to subagents and verifies its own work more readily than prior versions, and recommends institutions using it inside coding or research-automation harnesses set explicit, deterministic caps on subagent spawning and total spend rather than relying on the model to self-limit. That is a direct illustration of a broader point relevant to research-integrity and IT-governance offices: as models are configured to run longer, more autonomous, higher-effort chains of reasoning and tool use with less per-step human review, the operational question shifts from “is the output accurate” to “who set the effort and autonomy level this task ran at, and was that appropriate for what was being asked of it.” A literature-review synthesis run at low effort with light human review and one run at max effort as an unsupervised multi-step agent are not the same use of the tool, even when the underlying model and license are identical — and most current institutional AI-use policies do not yet distinguish between them.
Implications for AI-disclosure practice
Effort tiers also sharpen a gap in how AI-disclosure statements are typically written. A disclosure that names the tool and model version (“drafted with assistance from Claude Opus 5”) says little about how much independent reasoning the tool did versus how much a researcher directed and checked each step — a distinction that effort-tier selection now makes explicit and, in principle, auditable at the API level, even though almost no current journal or funder disclosure template asks for it. Research-integrity and publishing-ethics offices drafting or revising AI-use policy language have an opening here: disclosure guidance that asks not just which tool was used but what role it played (drafting, editing, literature retrieval, code generation, autonomous multi-step analysis) will age better than guidance tied to a specific model name, precisely because tier and mode, not model identity alone, are now doing much of the work of determining how autonomously the tool operated.
The pace problem
There is also a plainer, more mundane implication for institutions writing AI-use policy: the underlying technology is now changing meaningfully inside a single named model release, not just between releases. A policy written around “Claude Opus 5” as a fixed reference point already undersells what that name can mean in practice, since the same model can be configured across a real capability range depending on effort setting, and Anthropic has signaled it expects institutions to actively tune that setting against their own evaluations rather than treat a single default as correct for all uses. Static, annually-reviewed AI-tool policies are a poor fit for a landscape where a single model update changes the operationally relevant unit from “which model is licensed” to “which model, at which effort tier, for which task class.” Institutions maintaining AI-tool governance documents should expect to revisit configuration-level guidance — not just vendor and model approval lists — on a shorter cycle than has typically applied to policy review.
What to watch
- Whether vendors building research tools on top of Claude Opus 5 (reference managers, writing assistants, literature-synthesis tools, coding copilots for research software) disclose their default effort tier and make it configurable or auditable by institutional licensees.
- Whether journal, funder, or institutional AI-disclosure templates begin asking about the degree of autonomy or reasoning depth a tool was run at, not just which tool and model were used.
- Whether other frontier labs converge on a similar per-query effort-tier model — if so, “which effort tier, not just which model” becomes a durable procurement and governance question rather than an Anthropic-specific one.








