Written and maintained by CASRAI Editorial Board
Last updated
“AI compliance tool” is a marketing phrase before it is a product category. Five vendors can use it truthfully about five products that share almost nothing: one watches the Federal Register, one answers policy questions, one runs a conflict-of-interest disclosure queue, one tracks who has completed which training, and one reads documents looking for problems. They are procured differently, they fail differently, and they belong to different owners inside the institution.
Most bad purchases start with that confusion. This guide is for whoever runs the evaluation: what the categories actually are, which criteria predict whether a compliance office will trust the output, what to put to a vendor in writing, and which procurement and data-protection questions to settle before a demo is worth booking.
The five things sold as an “AI compliance tool”
Sort every candidate into one of these before comparing anything else. A product claiming all five is usually strong at one.
| Category | What it does | Natural owner | Failure mode |
|---|---|---|---|
| Regulatory change monitoring | Watches rulemaking, agency notices and funder policy pages; alerts on changes relevant to your portfolio | Research compliance / sponsored programs | Alert fatigue: high recall, no relevance filtering, nobody reads it by month three |
| Policy question answering | Retrieves and summarises an answer to “what does the rule say about X”, with sources | Compliance office; front-line administrators | Confident wrong answers outside its corpus |
| Disclosure and COI workflow | Collects disclosures, routes review, records management plans; AI assists triage and duplicate detection | COI office, often IT-owned as a system of record | Bought as an AI product, discovered to be a two-year workflow migration |
| Training and attestation tracking | Assigns, reminds, records completion and expiry against requirement rules | Research education / compliance operations | Rules engine cannot express your actual requirement matrix |
| Document review and screening | Reads protocols, subawards or agreements and flags issues against a checklist | Whichever office owns the document | Flags everything, so reviewers stop reading flags |
The structural split matters more than the feature list. Categories three and four are systems of record: they hold institutional data and need migration plans, HR and grants integration, and a retention policy. Categories one, two and five are systems of reference: they inform a human decision and hold little of your data. The first kind is a multi-year commitment; the second can be piloted and dropped. Do not run one procurement process for both.
Criteria that predict whether the compliance office will trust it
Demos are optimised for the questions vendors know they answer well. The criteria below separate a tool that gets used from one that dies quietly after the pilot, and each is testable in an hour.
Citation and provenance
Every factual claim should be traceable to a specific, retrievable source — not a bibliography of plausible references, but a link to the passage the claim came from, with a date. A compliance office cannot put an uncited summary into a determination or a response to a sponsor. Test it by clicking through: open three citations from three answers and confirm the cited passage actually supports the sentence attached to it. Fabricated or loosely-attached citations are a well-documented failure of generative systems — see hallucination and choosing and governing large language models for research.
Refusal behaviour
This is the strongest single predictor, and almost nobody tests it. Ask ten questions the tool should not be able to answer: your own local policy, a funder it does not cover, a rule that changed last week, a question with a false premise. A tool that answers all ten is dangerous. One that answers six and says “I don’t have that” for four is usable, because the boundary of its knowledge is visible to the person relying on it. An office burned once by a confident wrong answer will not come back, whatever the accuracy statistics say.
Corpus scope and currency
Ask what is in the corpus and how often it is re-ingested, then ask the harder question — what is not in it. US federal coverage is usually the strongest part of any product here; UK, EU and other national funders, state law and professional-society policy are usually much thinner. Get the source inventory in writing and attach it to the contract. “Real-time” and “always current” are marketing; an ingest cadence with a stated lag is the honest and more trustworthy version.
Reproducibility and audit trail
Ask the same question twice, an hour apart, and compare. Generative systems are not deterministic by default, so if an answer will be pasted into a record the office needs to know whether it can be reproduced later: are model versions pinned, are prior answers retrievable verbatim, what happens to an old answer when its source changes? That is the same concern driving data provenance practice generally. Alongside it you need a durable, exportable record of who asked what, what came back, which sources were cited and when. If the tool touches a regulated process, hold it to the same electronic-records expectations as anything else in scope — our audit trail review procedure and 21 CFR Part 11 compliance checklist show what a defensible trail looks like in a GxP context.
Data residency, retention and reuse
Three questions, in writing: where are prompts and uploaded documents processed and stored; how long are they retained; and are they used to train or improve any model. The third is where deals fall apart late, because the answer is often “not by us, but by our model provider unless you are on the enterprise tier”. Get the subprocessor list. If the tool will ever see protected health information, the BAA question is a gate rather than a negotiation point — see using AI with PHI in research and getting a BAA for AI tools used in research.
The human decision boundary
A tool should inform a determination and never make one. Ask where a human sign-off is recorded in the interface, and what happens when the reviewer disagrees with the AI. If disagreement is not modelled at all, the product was designed for a market where the machine decides — which is not research compliance.
What to ask the vendor
Send these before the demo and ask for written answers. The quality and speed of the reply tells you as much as the demo does.
Grounding and accuracy. Is the system retrieval-based, or relying on a model’s internal knowledge? If a vendor says the model is “trained on regulations”, ask specifically whether that means fine-tuning or retrieval-augmented generation — the implications for currency and citation are entirely different. Then: the full source inventory and per-source ingest cadence; a live demonstration of what happens to an out-of-corpus question; how accuracy is measured and on what evaluation set; whether model versions are pinned and whether you are notified before they change.
Security and data. Which subprocessors see your prompts and documents, in which jurisdictions; retention period and deletion mechanism; whether your content is excluded from model training by contract rather than by policy page; a current SOC 2 Type II report or equivalent (our note on SOC 2 compliance cost explains what sits behind one); SSO, role-based access, and audit logs exportable via API.
Commercial and exit. What is the pricing metric — seats, questions, documents, or institution-wide? Metered question pricing changes behaviour: people stop asking. What do you get on termination, in what format, and how long do you have to retrieve it? Do you have contractual audit rights over the security controls the vendor is asserting? And: three reference customers of your size and type, spoken to without the vendor present.
Procurement and data-protection questions to settle first
Each of these delays a purchase by weeks if raised late and costs nothing if raised first.
Federal award procurement standards
If any part of the cost sits on a federal award, the Uniform Guidance procurement standards apply to buying the tool itself. 2 CFR 200.318 requires that the recipient “must maintain and use documented procedures for procurement transactions under a Federal award or subaward”, that it “must maintain written standards of conduct covering conflicts of interest and governing the actions of its employees engaged in the selection, award, and administration of contracts”, and that it “must maintain records sufficient to detail the history of each procurement transaction”, including “the rationale for the procurement method, contract type selection, contractor selection or rejection, and the basis for the contract price”.
So the evaluation you are running is itself a record you must keep: write the criteria down before the demos, score against them, retain the scoring. If a panel member has a financial relationship with a bidder, that is a procurement conflict under those same written standards — a separate test from the investigator-facing NIH financial conflict of interest regime. For structuring the solicitation, see what an RFP for research procurement should contain.
Data protection, controlled and export-sensitive material
If personal data of people in the UK or EU will be processed — disclosure records, training records and anything a document-review tool ingests all qualify — you are likely looking at a data protection impact assessment before deployment, not after. Involve the data protection officer during shortlisting; a DPIA that surfaces an unacceptable transfer arrangement after signature is an expensive way to learn the answer. See GDPR and data protection compliance in research for the underlying obligations.
If the tool will touch controlled unclassified information, the hosting environment must meet the relevant safeguarding requirements before anyone uploads a document — see NIST SP 800-171 and CUI in university research. The same caution applies to export-controlled technical data: putting it into a commercial AI service with foreign-national support staff is a deemed-export question, not an IT question. Institutions building out NSPM-33 research security program obligations should route this through research security rather than treat it as routine software procurement.
Risk framing
Where the institution needs a defensible governance structure rather than an ad-hoc one, the NIST AI Risk Management Framework (version 1.0, released 26 January 2023) is the common reference point, organised around four core functions — Govern, Map, Measure and Manage. It is voluntary and not research-specific, but it gives a committee shared vocabulary and a structure for documenting why a tool was accepted.
Designing a pilot that tells you something
A two-week trial where three people have a play produces an opinion, not evidence. A pilot that produces evidence looks like this:
- Build the question set first. Take 30 to 50 real questions from your own ticket queue where you already know the correct answer and can prove it.
- Seed it with traps. Include at least five questions the tool should refuse: local institutional policy, an uncovered funder, a false premise, a question requiring a legal determination.
- Score in four buckets, not two: correct with a working citation; correct but uncited; wrong; appropriately refused. Refusals score positively.
- Have the actual user score it — the analyst who answers these questions today, not the person who championed the purchase.
- Time it. A tool that is 90% accurate but takes as long to verify as to answer from scratch has saved nothing.
Run the identical set against every shortlisted product and against a general-purpose assistant with no compliance specialisation as a control. If the specialised tool cannot beat the control on your own questions, the specialisation is not real.
Where these tools legitimately fail
- No tool makes you compliant. It surfaces a rule faster. The determination, documentation and liability stay with the institution, and a vendor claiming otherwise has told you something useful about the rest of their claims.
- Your local policy is in nobody’s corpus. The gap between what the federal rule says and what your institution requires is where most real questions live. Ask whether you can add your own policy documents, and what that costs.
- Non-US coverage is usually shallow. Test this specifically if you hold UK, EU, Canadian or Australian funding; a logo grid is not evidence of depth.
- Workflow products with AI bolted on are still workflow products. Evaluate a COI or training system on workflow, reporting and integration first — compare how narrow a well-scoped tool like electronic signature software can usefully be.
- Training platforms have a different job. Content coverage and requirement mapping matter more than intelligence; what CITI Program covers and does not satisfy is a useful reference when scoping that category.
Disclosure: one of these tools is ours
Ask CASRAI is our own product, so treat this as disclosure rather than a recommendation, and apply the tests above to it exactly as you would to any vendor’s.
It sits in the second category — policy question answering. It is a retrieval-based system over roughly 72,264 indexed passages from CASRAI’s own published pages, plus six external feeds: the US Federal Register, a research-and-grants slice of it, Grants.gov, Regulations.gov, NSF News and UKRI, re-ingested daily by a scheduled job. It cites a source for every factual claim and answers “I don’t have that in the CASRAI corpus” rather than guessing. It is a paid subscription. It answers questions; it does not draft documents. It does not guarantee compliance and does not cover every funder or regulator — the non-US caveat above applies to it too. If what you need is a disclosure workflow, a training tracker or a document-review engine, it is the wrong category of product and you should be looking at categories three, four and five.
Frequently asked questions
What is an AI compliance tool?
In a research setting it is any of five distinct products — regulatory change monitoring, policy question answering, disclosure and COI workflow, training and attestation tracking, or document review — using machine learning somewhere in the loop. The term names a marketing category, not a function; establish which of the five a product actually is before comparing prices.
Do AI compliance tools guarantee compliance?
No, and a vendor who implies otherwise should be disqualified on that basis. They reduce the time to find and interpret a requirement. The determination, the record and the consequences of getting it wrong remain with the institution and the responsible official.
Is a general-purpose AI assistant good enough?
Sometimes — test the assumption by running one as the control arm of your pilot. What justifies a specialised product is citation to retrievable sources, a defined and disclosed corpus with a stated update cadence, refusal outside that corpus, and an exportable audit trail. Where a general assistant already meets the need, the honest answer is that you need not buy anything.
What is the difference between a “trained on regulations” model and a retrieval system?
A fine-tuned model has absorbed material into its weights: it cannot cite a source and cannot be updated without retraining. A retrieval system looks up passages at query time and shows what it found, so it can cite, and it updates when the corpus updates. For compliance work retrieval is almost always the right architecture — ask which a vendor actually uses, because “trained on” gets applied loosely to both.
Can we build regulatory change monitoring ourselves?
Partly. Core US federal sources — the Federal Register, Regulations.gov, Grants.gov — publish machine-readable interfaces a competent developer can build alerting on. What a vendor adds is relevance filtering, coverage of sources with no clean feed (funder policy pages, society guidance, non-US regulators), and someone whose job it is to notice when a feed breaks. Price the maintenance, not just the build.
Who should own the evaluation?
The office that will live with the output, with IT security and the data protection officer as gates rather than decision-makers, and procurement involved from the start if federal funds are in play. An IT-owned evaluation over-weights integration and under-weights answer quality; a compliance-owned one tends to find the security problems after signature.
Related reading
For the wider category see our research tools and software hub; for governing general-purpose models rather than buying specialised ones, see choosing and governing LLMs for research.








