Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & Research SupplyReagents, PPE & instruments — chain-of-custody documented.Fast, traceable sourcing built for regulated research environments, from bench consumables to instrumentation.Shop lac.us CodeCASRAIlac.us

Editorial · CASRAI · AI and ML research outputs

Falling AI Token Costs and the New Math of Institutional Budgets

Frontier-model pricing now spans roughly two to three orders of magnitude. What that spread means for how research offices should structure AI-tool budgets and disclosure policy.

Published 16 Aug 2026· 6 minute read

Ask about this story

Answers are drawn from this article and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

CASRAI is the reference for research administration — bookmark it for the next question.

TL;DR: Current-generation AI model pricing spans an unusually wide range. Independent tracker Artificial Analysis lists some models, such as GPT-5.6 Luna Low and MiMo-V2.5, priced around $0.01 per million tokens blended, while a flagship reasoning model like Anthropic’s Claude Opus 5 holds a flat $5 per million input / $25 per million output tokens — and Google’s newly released Gemini 3.7 Flash sits in between at $0.75 input / $3.75 output. Those aren’t directly comparable numbers — blended pricing and separate input/output rates measure different things, and a $0.01 model and a $25 model are not interchangeable for the same task — but the spread is real and it is reshaping how research offices should think about AI as a budget line item: not as a single annual license fee, but as a tiered set of costs that should track task risk and stakes, not vendor brand.

What the current spread actually looks like

Three recent, independently verifiable data points illustrate the range research offices are now budgeting against:

  • Cheapest tier: Artificial Analysis’s own tracking (checked in the days around this article’s publication) puts the least expensive models it measures, including GPT-5.6 Luna Low and MiMo-V2.5, at roughly $0.01 per million tokens on a blended input/output basis — pricing aimed at high-volume, low-latency, agentic workloads rather than maximum reasoning depth.
  • Mid-tier: Google’s Gemini 3.7 Flash, released August 13, 2026, prices at $0.75 per million input tokens and $3.75 per million output tokens. On Artificial Analysis’s Intelligence Index it ranks 17th of 188 models tracked, but it is the fastest model the tracker currently measures (340 output tokens/second) — see CASRAI’s own coverage in Gemini 3.7 Flash and the Speed-Cost Case for Institutional AI Tools.
  • Frontier reasoning tier: Anthropic’s Claude Opus 5, released July 24, 2026, holds a flat $5 per million input tokens and $25 per million output tokens — unchanged from its Opus 4.8 predecessor even after Anthropic added configurable “effort tiers” that let a task consume more or less reasoning compute at the same per-token rate. Details in Claude Opus 5’s Adaptive Reasoning Tiers.

The gap between the cheapest and most expensive tiers here runs to roughly two to three orders of magnitude, depending on which numbers are compared. That is not a one-off; it reflects a deliberate market structure. Every major lab now ships both a cheap, fast, high-volume model and a separate, far pricier frontier reasoning model, and prices most of the total addressable market of routine tasks toward the cheap end while keeping frontier reasoning capacity expensive.

Why the gap exists, and why it is not closing

Vendors are not converging on a single price point because they are not competing for a single use case. Cheap, high-throughput models are priced for screening, triage, summarization-at-scale, and agentic pipelines where volume is high and any individual output is low-stakes and easily checked. Frontier reasoning models are priced for tasks where depth, accuracy on hard problems, and the cost of a wrong answer justify a premium — complex synthesis, high-stakes drafting, or judgment calls that would otherwise require a skilled person’s time. Because the two tiers serve different jobs, cutting the price of one doesn’t put competitive pressure on the other. Research offices that budget for “AI tools” as one line item are implicitly averaging across two products with a 100x-plus cost difference and materially different appropriate use cases.

What this means for institutional AI budgets

A few practical shifts follow from a market that looks like this rather than one with a single, converging price:

Budget by task tier, not by tool or vendor

Instead of a single “AI software” line item, research offices are better served treating high-volume/low-stakes automation (compliance screening, application intake triage, literature or dataset scanning) and low-volume/high-stakes generative work (analysis, drafting, decision support) as separate budget lines with separate approval thresholds. The cost profiles genuinely don’t compare, and neither does the appropriate level of human oversight.

Avoid long, flat-rate commitments

Given how fast headline prices have moved this year alone — a new mid-tier model launching at roughly one-tenth the cost of the prior generation’s frontier tier is now a routine occurrence rather than an outlier — a multi-year fixed contract locked to current pricing is a real financial-planning risk. Shorter renewal cycles, usage-based components, or contract language that references independent benchmark trackers rather than a fixed dollar figure give a research office more room to capture falling costs instead of being locked out of them.

Track quality-adjusted cost, not sticker price

The cheapest model in a category is not automatically the right procurement choice for a given task; a model’s rank on a standardized benchmark battery (Artificial Analysis and similar independent trackers publish these on a rolling basis) relative to its price is the more useful number for a procurement decision than either price or benchmark rank alone. Re-checking that comparison at renewal, not just at initial purchase, matters given how quickly rankings change.

Revisit the budget more often than the policy

An annual AI-tools budget set against a market moving this fast will be stale within a quarter. Building in a lighter-weight quarterly check-in on pricing and available tiers, separate from the heavier annual policy review, keeps the budget realistic without requiring the full governance apparatus to reopen every time a vendor ships a price change.

The governance angle: cheaper models change who is watching

Falling per-token cost for high-volume tiers has a research-integrity dimension that is easy to miss in a purely financial read of this trend. When screening, triage, or first-pass analysis becomes cheap enough to run on every application, dataset, or manuscript that crosses a desk, it stops being a deliberate, occasional tool choice a researcher or administrator makes and starts becoming default background infrastructure. CASRAI’s guide on choosing and governing LLMs for research and the site’s AI-disclosure guidance both note that disclosure frameworks built around “did a person consciously open a chatbot” struggle with exactly this shift. A falling price floor is what makes always-on, high-volume automated use financially trivial for an institution to adopt at scale — which means budget decisions about which task tier to fund are, in practice, also governance decisions about how much AI-mediated review and screening is happening with limited human checkpoints. Research offices setting next year’s AI budget are well served treating the two conversations — what tier we can now afford to run at scale, and what oversight that scale requires — as the same conversation, not two separate ones handled by different committees on different timelines.

A practical checklist for research offices

  • Split the AI budget into at least two tiers matched to task stakes, not a single line item.
  • Reference independent benchmark-and-price trackers (e.g., Artificial Analysis) at renewal, not just at initial purchase, since rankings and prices both move quickly.
  • Prefer shorter commitment terms or usage-based pricing over multi-year flat licenses while the market is still this volatile.
  • Pair any budget decision that increases automated-screening volume with an explicit check against current AI-disclosure policy — more affordable high-volume use is also more oversight-relevant use.
  • Track CASRAI’s AI writing and research tools coverage and independent leaderboards rather than relying on a single vendor’s pricing page, which reflects that vendor’s incentives, not the comparative landscape.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →