Written and maintained by CASRAI Editorial Board
Last updated
The GAD-7 (Generalized Anxiety Disorder 7-item scale) is a seven-question self-report screener for generalized anxiety, scored 0–21. It was developed by Robert L. Spitzer, Kurt Kroenke, Janet B.W. Williams and Bernd Löwe, and published in Archives of Internal Medicine in 2006 as a brief measure for detecting probable generalized anxiety disorder in primary care. It sits in the same family of short, validated bedside and clinic instruments as the Braden Scale, the Morse Fall Scale and the RASS: one number, a set of items underneath it that matter as much as the total, and a set of cutoffs that were derived against a reference standard rather than chosen by convention. This page covers the instrument itself — the seven items, how the total is built, what the severity bands and the validated cutoff actually mean, how it is used in primary care and as a clinical-trial outcome measure, and where it stops being informative.
The seven items
Each item asks how often, over the last two weeks, the patient has been bothered by the following problems. The recall window is part of the instrument — changing it invalidates the cutoffs.
- Feeling nervous, anxious, or on edge
- Not being able to stop or control worrying
- Worrying too much about different things
- Trouble relaxing
- Being so restless that it is hard to sit still
- Becoming easily annoyed or irritable
- Feeling afraid, as if something awful might happen
The seven items map onto the cognitive, somatic and behavioural features of generalized anxiety: items 1–3 capture the worry and apprehension core, items 4–5 the restlessness and tension, and items 6–7 irritability and dread. All seven load onto a single factor in the published psychometric work, which is what justifies summing them into one total rather than reporting subscales — see CASRAI’s construct validity guide for why unidimensionality is the precondition for a meaningful sum score.
How the score is built
Each item is answered on a four-point frequency scale, scored 0 to 3:
- 0 — Not at all
- 1 — Several days
- 2 — More than half the days
- 3 — Nearly every day
Sum the seven item scores for a total ranging from 0 to 21. There is no weighting, no reverse-scored item and no algorithm — it is a plain sum, which is a large part of why it survives in busy clinics. This is a frequency-anchored ordinal scale rather than an agreement scale, but the same analytic cautions apply as for any summed ordinal instrument; see CASRAI’s Likert scale design and analysis guide.
The eighth question is not scored
The standard GAD-7 form carries an eighth question after the seven items: if you checked off any problems, how difficult have these problems made it for you to do your work, take care of things at home, or get along with other people? — answered not difficult at all, somewhat difficult, very difficult, or extremely difficult. This item does not contribute to the 0–21 total. It is a functional-impairment probe, and it is the single most commonly mis-scored part of the instrument. A chart audit that finds totals above 21 has almost always found a site that is summing eight items instead of seven.
Severity bands
The total maps to four conventional severity bands:
- 0–4 — minimal anxiety
- 5–9 — mild anxiety
- 10–14 — moderate anxiety
- 15–21 — severe anxiety
These bands are descriptive severity anchors for communicating and tracking a score. They are not diagnostic categories, and they are not the same thing as the validated screening cutoff below — a distinction worth being precise about in policy documents, because the two are routinely conflated.
The validated cutoff, and what it was validated against
In the original 2006 validation, a cutoff of 10 or above was identified as the optimal threshold for detecting probable generalized anxiety disorder, yielding a sensitivity of 89% and a specificity of 82% against a structured diagnostic interview by a mental-health professional as the reference standard.
Two things follow from that, and both matter operationally:
- The cutoff is a screening threshold, not a diagnosis. At 82% specificity, a meaningful share of patients who screen positive will not have the disorder on structured assessment. A positive GAD-7 is an indication for further evaluation, not an entry for a problem list.
- The reference standard was generalized anxiety disorder specifically. The instrument was derived for GAD. It has been reported to detect panic disorder, social anxiety disorder and post-traumatic stress disorder at above-chance rates, but with lower accuracy than for GAD, because those were not what the cutoff was optimised against. Treat it as a GAD screener that is somewhat sensitive to related conditions, not as a general anxiety-disorder classifier.
GAD-2: the two-item short form
The first two items alone (feeling nervous or on edge; not being able to stop or control worrying) form the GAD-2, scored 0–6, with a positive screen at 3 or above. It is designed for the situation where even seven items are too many — a triage or intake step, or an ultra-brief combined screen. The usual practice is to treat the GAD-2 as a first-stage filter and administer the full GAD-7 to anyone who screens positive, which preserves most of the detection while cutting the burden on the negative majority.
Use in primary care and the USPSTF recommendation
The GAD-7 is the most widely deployed anxiety screener in United States primary care, and its position was formalised in June 2023, when the U.S. Preventive Services Task Force issued a Grade B recommendation to screen for anxiety disorders in adults 64 years or younger, including pregnant and postpartum persons. The Task Force identified the GAD scale versions — GAD-2 and GAD-7 — as the most commonly studied instruments in the evidence base it reviewed.
Two boundaries of that recommendation are easy to lose in implementation:
- It is age-bounded. For adults 65 and older, the Task Force issued an I statement: the current evidence is insufficient to assess the balance of benefits and harms of screening for anxiety disorders in older adults. That is not a recommendation against screening — it is an absence of evidence — but a hospital or health-system policy that asserts a universal screening mandate across all adult ages is claiming more guideline support than exists. The Task Force separately noted geriatric-specific instruments (the Geriatric Anxiety Scale and Geriatric Anxiety Inventory) in this population.
- A Grade B screening recommendation presumes adequate follow-up. Screening only produces benefit where there is capacity to evaluate and treat positives. Deploying the instrument without a defined downstream pathway converts a quality initiative into an unactioned-result liability — the same failure mode quality teams already track for critical results and early warning score escalation.
Use as a clinical-trial outcome measure
Beyond screening, the GAD-7 is used extensively as a continuous outcome measure — frequently the primary or a key secondary endpoint in trials of pharmacological, psychotherapeutic and digital interventions for anxiety. Its appeal as an endpoint is the same as its appeal in clinic: short, free, well characterised, and sensitive to change.
Three points bear on using it this way:
- Score it continuously, not by band. Collapsing a 0–21 score into four severity categories for analysis discards information and reduces power. Use the total as a continuous variable and report the bands descriptively.
- Change thresholds are conventions, not fixed constants. A change of roughly 4 points is the value most commonly cited in the literature as a minimally important difference on the GAD-7, but published estimates vary by population and by the anchor method used to derive them. If a protocol defines responder status by a point change, the specific threshold and its source should be pre-specified rather than assumed.
- Ceiling and floor behaviour constrain who can show change. A patient already at 0–2 cannot demonstrate improvement, and screening-enriched samples cluster differently from treatment-seeking ones. See CASRAI’s guide to floor and ceiling effects for how this attenuates measured effect sizes.
For the underlying measurement questions — reliability, validity evidence, and what a summed self-report score can support — see CASRAI’s psychometrics overview, Cronbach’s alpha guide, and convergent and discriminant validity guide.
Pairing the GAD-7 with the PHQ-9
The GAD-7 was built by the same research group as the PHQ-9, the nine-item depression measure scored 0–27, and the two share a response format, a two-week recall window and a scoring logic. That is a deliberate design property, not a coincidence, and it is why the pair is the default combination for joint depression-and-anxiety screening.
The clinical case for administering both is comorbidity: anxiety and depressive disorders co-occur at high rates, and either instrument alone will systematically miss patients whose predominant presentation falls on the other side. Screening for one where both are plausible produces a confidently negative result on a question that was not asked.
Three combined configurations are in common use:
- PHQ-9 plus GAD-7 administered together — 16 items, two separate totals, each interpreted against its own cutoff. Both scores are reported; they are not summed into a single number in routine practice.
- PHQ-4 — the ultra-brief four-item combined screener formed from the PHQ-2 and the GAD-2, scored 0–12. Used as a single first-stage filter, with the full PHQ-9 and GAD-7 administered when it is positive.
- Stepped screening — GAD-2 and PHQ-2 at intake, escalating to the full instruments on a positive. Operationally the same idea as the PHQ-4, expressed as two sequential steps.
One asymmetry is worth building into any combined workflow: the GAD-7 contains no suicide-risk item, and the PHQ-9 does (its ninth item asks about thoughts of being better off dead or of self-harm). A screening program that deploys the GAD-7 alone therefore has no suicidal-ideation trigger anywhere in it, and needs one from another source. This is a concrete patient-safety gap, not a documentation preference.
Limitations: what a screening instrument cannot do
The GAD-7 is a well-validated screener. It is not a diagnostic test, and the distance between those two things is where most misuse lives.
- It cannot diagnose. No GAD-7 score establishes generalized anxiety disorder. Diagnosis requires clinical assessment against DSM criteria, including duration, distress or functional impairment, and the exclusion of substance effects and other medical and psychiatric causes. A score of 15 is a strong indication to evaluate, not a diagnosis of severe GAD.
- It is self-report. It measures what a patient reports over the last two weeks, and is subject to the same recall, social-desirability and health-literacy effects as any self-administered questionnaire. It can be minimised by a patient who does not want the finding on their record, and inflated by acute situational distress that is not a disorder.
- Somatic items overlap with medical illness. Restlessness and irritability are not specific to anxiety; they occur in pain, delirium, thyroid disease, medication effects and withdrawal states. In an acutely ill inpatient population, the specificity established in ambulatory primary care should not be assumed to transfer.
- Positive predictive value depends on prevalence. Fixed sensitivity and specificity produce very different post-test probabilities across settings. Applied to a low-prevalence population, a cutoff of 10 generates a higher proportion of false positives than the same cutoff in a treatment-seeking population — a screening-program design question, not a flaw in the instrument.
- Evidence is thinner at the age boundaries. The USPSTF I statement for adults 65 and older reflects genuine uncertainty in this population, and paediatric use requires separately validated instruments and cutoffs rather than an assumption that the adult thresholds transfer.
- Translations are not interchangeable by default. Validated translations exist in many languages, but a locally produced or ad hoc translation is a different instrument until it has been validated, and should not inherit the original cutoffs.
Licensing and reproduction
The GAD-7 is free to use. The instrument was developed with an educational grant from Pfizer Inc., and the standard form carries an explicit statement that no permission is required to reproduce, translate, display or distribute it. This matters practically: unlike several proprietary assessment instruments, the GAD-7 can be embedded in an EHR flowsheet, printed on an intake packet or translated for a local population without a licence fee or a use agreement. Verify the wording on the specific copy of the form your organisation has adopted rather than relying on a general assumption of permissiveness.
Operational notes for quality and patient-safety staff
The recurring implementation failures with the GAD-7 are administrative rather than clinical, and each is auditable:
- Scoring the eighth item. The functional-impairment question is not part of the total. Any total above 21 is a scoring defect — and a build that permits it should be corrected, not just retrained around.
- Recording a band without a total. A charted band alone cannot be recalculated, audited or tracked as a continuous measure over time. Store the integer total.
- Altering the recall window. Changing "over the last two weeks" to a different period produces a locally invented instrument with borrowed cutoffs.
- Positives without a documented action. As with any risk score, what a surveyor looks for is that the score drove something. A charted score of 14 with no assessment, referral, or documented clinical decision is the same defect pattern as a high fall-risk score with no corresponding interventions charted.
- Missing items. An incomplete GAD-7 is not a lower score. If items are unanswered, the total is not comparable to a complete administration and should be flagged rather than summed as if the blanks were zeros.
For the coding and documentation side of behavioural-health encounters that follow a positive screen, see CASRAI’s guide on E/M with a psychotherapy add-on, and what is psychiatry for the wider clinical and research landscape.
Frequently asked questions
What is a normal GAD-7 score?
A total of 0–4 falls in the minimal-anxiety band and is the conventional normal range. Scores of 5–9 indicate mild anxiety, which typically warrants monitoring rather than immediate intervention.
What GAD-7 score indicates anxiety?
A score of 10 or above is the validated screening cutoff for probable generalized anxiety disorder, with a reported sensitivity of 89% and specificity of 82% against a structured diagnostic interview. It indicates the need for further clinical evaluation, not a diagnosis on its own.
How is the GAD-7 scored?
Score each of the seven items 0 (not at all), 1 (several days), 2 (more than half the days) or 3 (nearly every day), then sum them for a total of 0 to 21. The eighth functional-impairment question on the standard form is not included in the total.
What is the difference between the GAD-7 and the GAD-2?
The GAD-2 is the first two items of the GAD-7, scored 0–6, with a positive screen at 3 or above. It is used as an ultra-brief first-stage filter, with the full GAD-7 administered to those who screen positive.
Should the GAD-7 be used with the PHQ-9?
In most screening programs, yes. Anxiety and depression are highly comorbid, the two instruments share a response format and recall window, and either alone will miss patients presenting predominantly with the other condition. Note that only the PHQ-9 contains a suicide-risk item.
Can the GAD-7 diagnose generalized anxiety disorder?
No. It is a screening and severity-tracking instrument. Diagnosis requires clinical assessment against DSM criteria, including duration, functional impairment, and exclusion of other medical, substance-related and psychiatric causes.
Is the GAD-7 free to use?
Yes. It was developed with an educational grant from Pfizer Inc., and the standard form states that no permission is required to reproduce, translate, display or distribute it.
How often should the GAD-7 be repeated?
The instrument does not specify a cadence. In practice it is used at screening intervals set by organisational policy and repeated to track response to treatment, commonly at each follow-up visit during active management. The reassessment schedule is a policy decision that should be documented and defensible, not an instrument property.
Back to the CASRAI Patient Safety hub for surveillance definitions, root cause analysis, risk instruments, and the rest of the hospital patient-safety and quality library.








