Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & Research SupplyReagents, PPE & instruments — chain-of-custody documented.Fast, traceable sourcing built for regulated research environments, from bench consumables to instrumentation.Shop lac.us CodeCASRAIlac.us

Choosing and Governing Large Language Models for Research: A Framework

General AI leaderboards rank intelligence, speed, cost, and context window — not the things that decide whether a model is safe for research work: citation fabrication, reproducibility, confidentiality, and disclosure obligations. This guide sets out the research-specific evaluation criteria and points to where live leaderboard data actually lives.

Ask about Choosing and Governing Large Language Models for Research: A Framework

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

General AI leaderboards rank models on intelligence, speed, cost and context window. None of that tells a research office, principal investigator, librarian or research-integrity officer whether a given model is safe to use on research work — because the failure modes that matter in research (fabricated citations, non-reproducible outputs, leaked unpublished data, undisclosed authorship) are not what those leaderboards measure. This guide sets out the evaluation criteria that are specific to research use, points to where the general leaderboards actually live so you can check current standings yourself, and covers the disclosure and confidentiality obligations that apply regardless of which model you choose.

What general leaderboards measure — and what they don’t

General-purpose LLM leaderboards are a genuinely useful, continuously updated signal for raw capability — but they are built to answer “which model is smartest, fastest, or cheapest,” not “which model is safe to cite in a manuscript.” Treat them as one input to a research-specific decision, not the decision itself.

Artificial Analysis

Artificial Analysis (artificialanalysis.ai) publishes a live model leaderboard scored across four axes: an aggregate “Intelligence Index,” output speed (tokens/second) and time-to-first-token latency, cost per task, and maximum context window. As of 14 August 2026, per a direct check of its leaderboard page, it does not publish a factual-accuracy, citation-accuracy, or hallucination-rate metric as part of its standard scoring. Standings on sites like this change frequently as providers ship new model versions — check the live page rather than relying on a number frozen into any article, including this one.

LMArena / Arena

LMArena — now operating as Arena (its leaderboard has redirected to arena.ai) — ranks models by blind, pairwise human-preference voting: real users compare two anonymized model outputs side by side and vote for the better response, and an Elo-style score is computed from the aggregated votes. This measures which output people prefer, across categories such as text, vision, and document tasks — not whether either output was factually accurate or its citations real. A model can win a preference vote for a fluent, well-formatted answer that is confidently wrong.

Why neither is a substitute for a research-specific check

Preference voting and capability indices both reward fluency, helpfulness, and task completion. A research administrator’s actual risk surface — a fabricated DOI in a literature review, an unrepeatable methods-section output, unpublished data pasted into a third-party tool, an undisclosed AI contribution flagged post-publication — sits outside what either measures. The rest of this guide covers that surface directly.

Citation and reference fabrication (“hallucinated” citations)

The single most reported research-specific failure mode of general-purpose LLMs is confident fabrication of references: plausible-looking author names, journal titles, volume/page numbers, and DOIs that do not correspond to any real publication, generated with the same fluency and formatting confidence as a genuine citation. This happens because an LLM predicts a statistically plausible continuation of text, not because it looks anything up by default — a model with no live retrieval or tool-use step has no mechanism to confirm a reference exists.

Practical mitigation, regardless of which model is in use:

  • Treat every citation an LLM produces as unverified until you have independently resolved it — search the DOI on Crossref or the title in a real database, don’t just check that the DOI “looks” well-formed.
  • Prefer models or tool configurations that ground answers in retrieval against a real corpus (search-augmented modes, connected reference-manager integrations) over a base chat model with no retrieval step, when the task is literature-facing.
  • Never paste an LLM-generated reference list directly into a manuscript’s bibliography without resolving every entry against a primary source first.

Reproducibility of outputs

The same prompt sent to the same model can return materially different text on repeat runs, because most LLM inference involves a sampling step (temperature/top-p) that introduces randomness by design, and because commercial providers periodically update the underlying model version without necessarily changing its public name. Neither general leaderboard covers this dimension.

This matters specifically for a methods section: if AI assistance shaped how an analysis script, a search strategy, or a data-cleaning step was written, “I asked the model” is not a reproducible method description in the way a fixed software version and parameters are. Where AI-generated code or text materially shaped a reported method, document the model name, version/date, and — where feasible — the prompt or a saved transcript, rather than describing the step only in terms of the model’s output. Some providers offer a deterministic or low-temperature mode; using it for research-facing tasks, and recording that you did, is a defensible practice even though it doesn’t eliminate version drift over time.

Confidentiality: what must never go into a third-party model

Pasting text into a hosted, third-party LLM is functionally similar to sending it to an external service you don’t control — unless a specific enterprise agreement states otherwise, assume the vendor may log, retain, or use the input in ways your institution has not reviewed. This is a data-governance question before it is an AI-quality question.

Content that should not be pasted into a consumer-grade or unreviewed LLM tool includes:

  • Unpublished data containing identifiable participant information — this is an IRB/human-subjects-protections issue independent of the consent form’s wording, since the vendor is a new, undisclosed data recipient.
  • Material under active peer review — most journals treat manuscripts and reviewer comments under review as confidential; submitting either to a third-party AI tool can itself be a breach of that confidentiality, separate from any question about AI-assisted review.
  • Data covered by a data use agreement, an NDA, or export-control restrictions — these typically name the permitted recipients explicitly, and a general-purpose AI vendor is not one of them by default.
  • Pre-publication or embargoed findings where confidentiality is contractual (e.g., under a funder or publisher embargo).

Where an institution has a vetted enterprise deployment (with a data-processing agreement, no default training on inputs, and defined retention terms), that changes the calculus for some of the above — check the specific agreement rather than assuming any “enterprise” or “business” tier is automatically safe. See CASRAI’s guidance on data management plans for how confidentiality commitments are typically documented before data collection even begins.

Disclosure obligations: what journals and funders require

Disclosure requirements for AI-assisted research output are now a settled, if still-evolving, area of publisher and funder policy — but the specifics vary by venue, and checking the actual policy of the specific journal or funder you’re submitting to remains necessary before you assume a general rule applies.

Two points are broadly consistent across the major publishing-ethics bodies:

  • An AI tool cannot be listed as an author. COPE’s position statement on authorship and AI tools states that AI tools cannot be listed as authors because they cannot take responsibility for the work and, as non-legal entities, cannot hold copyright, manage a license agreement, or declare a conflict of interest. See CASRAI’s dictionary entries on the ICMJE’s 2023 rejection of AI co-authorship for the equivalent position from the medical-journal editors’ body.
  • Use of generative AI in writing, image/graphical-element production, or data analysis must be disclosed — typically in the Methods/Acknowledgments section, naming the tool and describing how it was used, not merely that AI was “involved.”

What counts as an acceptable disclosure statement, and where it belongs in the manuscript, differs by publisher — CASRAI’s generative-AI disclosure statement entry and the guide to journal and publisher policies on generative AI in manuscripts track that variation directly. If your work uses the CRediT taxonomy to document contributions, see the guide on how to disclose AI assistance in a CRediT statement, role by role — CRediT roles are defined for human contributors, and mapping AI assistance onto them correctly (rather than awkwardly listing a tool as if it were a contributor) is a genuinely common point of confusion. For the broader legal-versus-editorial-policy distinction (jurisdictions increasingly have their own AI-disclosure statutes on top of publisher rules), see AI disclosure laws: the legal landscape vs. publisher policy.

Provenance and auditability of AI-assisted outputs

Provenance, in this context, means being able to reconstruct after the fact what an AI tool contributed to a given output, when, and using what input — the same audit trail an integrity investigation or a data-sharing request would expect for any other research record. Most consumer chat interfaces do not preserve this by default in a form suitable for a research record: conversation history can be edited, deleted, or tied to a personal account rather than a project.

Practical steps that make AI assistance auditable rather than just “used”:

  • Record the model name and version/date, not just “ChatGPT” or “an AI tool” generically — version drift is real and a vague reference is not verifiable later.
  • Save the actual prompt and output for any AI-assisted step that materially shaped a reported result, analysis, or figure, in a project-linked location rather than a personal chat history.
  • Where an institution or lab maintains an electronic lab notebook or project log, treat a substantive AI-assisted step the same way you’d log any other methodological decision — see CASRAI’s coverage of research tools for how lab informatics and ELN systems fit into that record-keeping.

CASRAI’s AI provenance entry covers this concept in more depth.

How to check a journal’s or funder’s AI policy before you use a model

Do this before, not after, you use AI assistance on a submission-bound piece of work:

  1. Check the specific journal’s “Instructions for Authors” or editorial policy page for a generative-AI or “AI tools” section — don’t assume the publisher’s group-wide policy (e.g., a large publisher’s general stance) is identical to every individual journal’s, since some vary at the title level.
  2. Check the funder’s grant terms and conditions, if the work is funded — several major funders have issued separate guidance on AI use in proposal preparation and in funded research output, distinct from journal policy.
  3. If either policy is silent or ambiguous, disclose anyway, in the Methods or Acknowledgments section, naming the tool and describing its role — disclosure is the lower-risk default when a policy doesn’t explicitly address your use case.
  4. Re-check before resubmission to a different venue if a manuscript is rejected and redirected — policies are not uniform across journals, even within the same publisher.

Institutional AI acceptable-use policy

Many of the choices above — which models are approved for institutional use, what may and may not be pasted into them, and how AI-assisted work must be documented — are exactly what an institutional AI acceptable-use policy exists to fix in advance, rather than leaving each researcher to decide case by case. See CASRAI’s guide to writing an AI acceptable-use policy for the clauses, scope, and enforcement mechanisms such a policy typically covers.

A note on AI writing tools and this site’s affiliate relationships

CASRAI’s AI writing tools hub covers specific commercial tools researchers use for drafting, editing, and citation support. CASRAI has affiliate relationships with some of the tools covered there — see the affiliate disclosure page for the full terms. This guide does not rank or recommend specific tools; it sets out criteria so you can evaluate any tool, including ones covered by that hub, against your own institution’s requirements.

Separately, CASRAI operates its own credential in this area — AI in Research Governance (AIRF/AIRP) — described on the certifications page. This is disclosed here plainly as CASRAI’s own credential, not an independent third-party standard, and enrolment is not currently open.

A practical checklist before adopting a model for research work

  • Confirm whether the deployment is a consumer tier, a paid consumer tier, or a vetted enterprise agreement with a data-processing agreement — this determines what data can safely go into it.
  • Confirm whether the target journal or funder has a published AI policy, and read it before drafting begins, not after.
  • Decide in advance how AI-assisted steps will be logged (model, version, date, prompt/output retained) so disclosure and reproducibility aren’t reconstructed from memory later.
  • Treat any citation, reference, or factual claim the model produces as unverified until independently checked against a primary source.
  • Check the live leaderboards (Artificial Analysis, Arena) for current capability/cost trade-offs only after the above governance questions are settled — capability is the last filter, not the first.

Frequently asked questions

Can an AI tool be listed as an author on a manuscript?

No. Both COPE’s position statement on authorship and AI tools and the ICMJE’s guidance reject AI co-authorship, on the reasoning that an AI tool cannot take responsibility for the work and, as a non-legal entity, cannot hold copyright or declare a conflict of interest. See CASRAI’s entry on the ICMJE’s 2023 position.

Is it safe to paste unpublished data into ChatGPT, Claude, Gemini, or a similar tool to help analyze it?

Not by default. Unless your institution has a specific vetted enterprise agreement covering that tool — with a data-processing agreement and defined retention/training terms — assume a consumer or standard paid tier may log or retain what you paste, and treat unpublished, identifiable, or contractually restricted data as off-limits until you’ve confirmed otherwise with your institution’s data-governance or IT-security office.

Do I need to disclose AI assistance even if the AI didn’t write the final text?

Generally, yes, if it materially shaped the writing, an image or graphical element, or the data analysis — publisher and integrity-body guidance (COPE, and journal-specific policies) frame disclosure around whether AI meaningfully contributed to the reported work, not only whether its raw output appears verbatim in the final manuscript. Check the specific journal’s policy, since exact thresholds for what must be disclosed vary.

Which LLM should I use for my research?

This guide deliberately does not answer that with a single winner, because capability rankings change monthly and a static answer would go stale within weeks on a topic where accuracy matters. Use the criteria above — citation reliability, reproducibility needs, confidentiality of your inputs, and your target venue’s disclosure policy — to narrow the field, then check a current, live leaderboard (Artificial Analysis or Arena) for capability and cost trade-offs among the models that pass those filters.

Do general AI leaderboards measure hallucination or factual accuracy?

Not as a rule. As checked directly on 14 August 2026, Artificial Analysis’s standard leaderboard scores intelligence, speed, cost, and context window; Arena (formerly LMArena) ranks models by blind human-preference voting, which measures which response people prefer, not whether it’s factually correct or its citations are real. Neither is a substitute for independently verifying an LLM’s output in a research context.

What should a lab or department’s AI acceptable-use policy cover, at minimum?

At minimum: which tools/tiers are approved for institutional use, what categories of data may never be entered into an unapproved tool, how AI-assisted contributions to research output must be documented and disclosed, and who to contact with questions. See CASRAI’s guide on writing an AI acceptable-use policy for the fuller clause-by-clause structure.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →