On this page: what a “rater” is in a clinical trial, why Clinical Outcome Assessment (COA) raters need formal training and certification, what that training/certification process typically covers, the regulatory basis for it, and how it differs from the electronic Patient-Reported Outcome (ePRO) modality where no rater is involved.
What is a “rater” in a clinical trial?
A rater is a study-site staff member — often a physician, nurse, psychologist, or trained clinical research coordinator — who administers a clinical outcome assessment scale and records a judgment about the participant’s condition, or who scores a standardized performance task. Rater training and certification apply specifically to the two COA categories where a human being, rather than the patient, generates the observation:
- Clinician-Reported Outcome (ClinRO) — a trained health-care professional observes the patient and applies clinical judgment to rate a sign or symptom on a defined scale (for example, a depression severity scale or a disease-specific clinical rating instrument).
- Performance Outcome (PerfO) — the patient performs a standardized task (a timed walk, a grip-strength test, a cognitive task) and a trained assessor administers the task consistently and scores the result according to a fixed protocol.
This is a distinct methodological concern from ePRO and Patient-Reported Outcome (PRO) instruments generally, where the patient self-reports directly with no intermediate rater judgment — see the last section below for why that distinction matters operationally.
Why rater training and certification exist
Any assessment that depends on a person’s judgment is vulnerable to variability: two raters can watch the same patient and score the same scale differently, and the same rater can drift in how they apply a scale over the course of a long study. In a clinical trial, that variability is not a minor inconvenience — it adds noise directly into the endpoint the trial exists to measure, which can obscure a real treatment effect or, worse, manufacture an apparent one. Regulators treat this as a data-quality issue with the same seriousness as any other source of measurement error.
The FDA and the European Medicines Agency both recommend that clinical study personnel administering ClinRO and PerfO instruments be trained, and FDA good clinical practice expectations call for documentation of the qualification requirements for each scale used in a study, the training provided to each rater, and the certification result for every rater who participated. Where a scale is used repeatedly across a longer study, periodic retraining or recertification is commonly expected to guard against rater drift — gradual, uncorrected changes in how an individual rater applies a scale over time.
What rater training and certification typically cover
Programs vary by sponsor, instrument, and therapeutic area, but a rater training and certification program for a ClinRO or PerfO instrument generally includes:
- Standardized administration training — instruction on exactly how to administer the scale or task: what to say to the participant, what environmental conditions to control, how to sequence sub-items, and what NOT to do (leading questions, inconsistent prompting).
- Scoring calibration — training on how to translate an observation into the scale’s defined score categories, typically using worked examples or video vignettes with a known, agreed-upon “gold standard” score.
- Certification testing — the rater scores a set of standardized test cases (often video-recorded patient encounters) and their scores are compared statistically against the gold-standard score or against a panel of other trained raters before the rater is approved to rate live study participants.
- Inter-rater and intra-rater reliability statistics — agreement is typically quantified using measures such as Cohen’s kappa, an intraclass correlation coefficient (ICC), or Pearson’s correlation, depending on whether the scale produces categorical or continuous scores.
- Ongoing surveillance and recertification — for longer trials, sponsors commonly monitor rater performance over time and require periodic retraining or recertification to catch and correct rater drift before it affects the study’s endpoint data.
Regulatory and quality-system basis
ICH E6 good clinical practice (in both the long-standing E6(R2) form and the newer E6(R3) revision) requires that everyone involved in conducting a trial be qualified by education, training, and experience to perform their assigned tasks — a general principle that applies squarely to anyone administering a ClinRO or PerfO instrument. FDA’s Patient-Focused Drug Development (PFDD) guidance series, which sets out how sponsors should select, develop, or modify a COA to be fit for purpose, treats rater qualification and training as part of demonstrating that a COA-based endpoint is reliable evidence of a treatment effect, not just that the instrument itself is well designed.
In practice this means rater training and certification records — who was trained, on what version of the training, when, with what certification result, and any retraining history — form part of a trial’s essential documents and are a routine focus of monitoring visits and regulatory inspections. A well-validated COA instrument administered by untrained or uncertified raters is still a source of unreliable data from a reviewer’s perspective.
Who provides rater training and certification
Rater training can be run in-house by a sponsor or by the study’s CRA/clinical operations team, but for licensed, widely-used clinical rating scales it is frequently outsourced to specialized central rating and rater-training vendors, who build standardized e-learning modules, certification test batteries, and ongoing rater-surveillance programs that can be applied consistently across every site in a multi-site, multi-country trial. Centralizing training this way is one of the more effective ways sponsors have found to keep inter-rater reliability high across a large, geographically dispersed rater pool — a much harder problem to solve with site-by-site, ad hoc training alone.
The FDA Clinical Outcome Assessment Compendium
Sponsors selecting or designing a COA instrument — and, by extension, planning the rater training program that instrument will require — commonly start with FDA’s Clinical Outcome Assessment Compendium, which collates COAs FDA has seen used to support labeling claims across many disease areas, including which have been qualified under CDER’s Drug Development Tool program. Inclusion in the compendium is informational, not an endorsement that a listed COA is or should be the sole determinant of effectiveness in a new trial — sponsors still need to justify a specific instrument’s fitness for their own concept of interest and context of use, which includes planning how raters for that instrument will be trained and certified.
Rater training vs. ePRO: why the two are different problems
It is worth being precise about scope, because the terms get blurred in practice: rater training and certification is a workforce/measurement-quality concern that applies to ClinRO and PerfO instruments, where a human rater’s judgment sits between the patient and the recorded score. ePRO — the electronic capture of a Patient-Reported Outcome directly from the patient, without amendment or interpretation by clinical staff — removes that intermediate judgment entirely, which is precisely why ePRO doesn’t carry a rater-training requirement in the same sense. ePRO still requires site staff to be trained on device provisioning, compliance monitoring, and data handling, but that is device/process training, not rater calibration against a scoring standard. Sponsors running a trial with both a ClinRO/PerfO component and an ePRO component should expect two distinct training programs, not one.
Frequently asked questions
Does every COA require rater training?
No. Rater training and certification apply to ClinRO and PerfO instruments, where a clinician or trained assessor generates the score. Patient-Reported Outcomes (including ePRO) and most Observer-Reported Outcomes captured by an untrained caregiver do not involve a “rater” in this formal, certified sense, though caregivers reporting an ObsRO are still typically given instructions on how to complete the report consistently.
What is “rater drift” and why does it matter?
Rater drift is the gradual, usually unintentional change in how an individual rater applies a scale over the course of a study — becoming more lenient or more strict over time, for example. Left unchecked, drift introduces a systematic bias into the trial’s endpoint data that is difficult to distinguish from a genuine treatment effect, which is why longer trials commonly build in periodic recertification or ongoing reliability monitoring.
How is inter-rater reliability measured?
The specific statistic depends on the scale’s data type: Cohen’s kappa is common for categorical/ordinal ratings, an intraclass correlation coefficient (ICC) is common for continuous scores, and Pearson’s r is sometimes used as a simpler correlation measure. Sponsors typically specify the reliability threshold a certification test must meet before a rater is approved to rate live participants.
Is rater certification the same as investigator qualification under GCP?
They’re related but not identical. ICH GCP‘s general requirement that trial personnel be qualified by education, training, and experience is the broader principle that rater certification sits underneath — rater certification is the scale-specific, documented demonstration that a particular individual meets that qualification standard for a particular instrument, which is a narrower and more measurable requirement than general investigator qualification.
Related CASRAI terms and guides
- Clinical Outcome Assessment (COA) — the umbrella concept: PRO, ClinRO, ObsRO, and PerfO defined
- ePRO (Electronic Patient-Reported Outcomes)
- ICH E6(R3) (Good Clinical Practice Guideline Revision)
- ICH GCP (Good Clinical Practice)
- Good Clinical Practice (GCP) Certification: What It Is, Who Needs It, and How to Get It
- Clinical Research Coordinator (CRC)
- Clinical Research Associate (CRA)
- Clinical Research Administration (cluster overview)
References
- FDA, Clinical Outcome Assessment Compendium — fda.gov/drugs/development-resources/clinical-outcome-assessment-compendium
- FDA, Patient-Focused Drug Development Glossary — fda.gov/drugs/development-approval-process-drugs/patient-focused-drug-development-glossary
- FDA-NIH, BEST (Biomarkers, EndpointS, and other Tools) Resource, Glossary — ncbi.nlm.nih.gov/books/NBK338448/
- ICH, E6(R2)/E6(R3) Good Clinical Practice guideline — ich.org







