In measurement, accuracy and precision describe two different, independent kinds of error, and confusing them leads to trusting instruments and results that don’t deserve it. Accuracy is how close a measurement is to the true value. Precision is how close repeated measurements are to each other. A measurement can have either quality without the other, and the most dangerous case — precise but inaccurate — is exactly the one that looks most trustworthy at a glance.
The core distinction
The classic illustration is a dartboard. Accuracy is whether the darts cluster around the bullseye (the true value). Precision is whether the darts cluster tightly together at all, regardless of where on the board that cluster sits.
Translated into an actual measurement scenario: suppose a laboratory balance is used to weigh a calibrated 100 g reference mass five times.
- If the five readings are 99.9, 100.1, 100.0, 99.9, 100.1 g, the balance is both accurate (the average is very close to 100 g) and precise (the readings are tightly clustered).
- If the five readings are 98.1, 102.3, 99.4, 101.0, 97.8 g, the average may still land near 100 g, but the spread is wide — this is accurate on average but imprecise, and any single reading is unreliable.
- If the five readings are 102.0, 102.1, 102.0, 102.2, 102.1 g, the balance is highly precise — the readings barely vary — but every one of them is off by about 2 g. This is a precise but inaccurate instrument, and it is the case that fools people, because tight, repeatable numbers look like reliable numbers.
Accuracy error is systematic: it shifts every reading in a consistent direction (a miscalibrated scale, a biased survey instrument, a thermometer that reads consistently high). Precision error is random: it scatters readings unpredictably around whatever their central tendency happens to be (electrical noise, small sampling variation, operator handling differences).
All four combinations
| Combination | What it looks like | Practical risk |
|---|---|---|
| Accurate and precise | Readings cluster tightly around the true value | The goal — low risk, both systematic and random error are controlled |
| Accurate but imprecise | Readings scatter widely but average out near the true value | Individual readings can’t be trusted; only the average is useful, and only with a large enough sample |
| Precise but inaccurate | Readings cluster tightly but around the wrong value | The dangerous case — low variability creates false confidence that masks a consistent bias |
| Inaccurate and imprecise | Readings scatter widely and don’t average out near the true value | Obviously unreliable, but at least it’s visibly unreliable, so it’s harder to trust by mistake |
The “precise but inaccurate” case is worth dwelling on because it is the one that survives casual quality checks. A researcher who only looks at variability between runs (precision) and never checks against an independent reference value (accuracy) can validate an instrument or a coding scheme that is quietly, consistently wrong.
How this maps onto reliability and validity
Research-methods vocabulary and metrology vocabulary describe the same two ideas with different words, and the mapping is close enough to be genuinely useful:
- Precision ≈ reliability. Both describe consistency — whether repeated application of the same measurement procedure gives the same result. See reliability in research measurement for how this is assessed with test-retest reliability, internal consistency (Cronbach’s alpha), and inter-rater agreement.
- Accuracy ≈ validity. Both describe correctness — whether a measurement actually reflects the true value or construct it’s supposed to capture. See types of validity in research for criterion, construct, content, and face validity.
This mapping is not exact — validity is a broader concept than accuracy alone, since it also covers whether an instrument measures the right construct at all, not just whether it measures a known quantity correctly — but for numeric measurement it is close enough to be the practical bridge between the two vocabularies. A researcher who has established high reliability (precision) for an instrument still has not established validity (accuracy); the two claims require separate evidence.
Quantifying accuracy
Accuracy is quantified relative to a known or accepted reference value (also called the true value or, in metrology, the conventional true value):
- Bias / systematic error — the average difference between measured values and the reference value. A balance that reads consistently 2 g high has a bias of +2 g.
- Trueness — the metrology term (ISO 5725) for the closeness of agreement between the average of a large number of measurements and the true value; it is essentially the inverse of bias.
- Percent error — bias expressed as a proportion of the reference value, useful for comparing accuracy across instruments with different scales.
Establishing accuracy requires an independent reference standard. Precision alone, no matter how tight, never tells you whether that reference has been met.
Quantifying precision
Precision is quantified from the spread of repeated measurements around their own mean, with no reference value required:
- Standard deviation (SD) — the basic measure of spread; see descriptive statistics for how it’s calculated.
- Coefficient of variation (CV) — SD expressed as a percentage of the mean, which makes precision comparable across instruments or scales with different units. Covered in depth in the coefficient of variation guide.
- Repeatability and reproducibility coefficients, often assessed formally with the intraclass correlation coefficient (ICC) when multiple raters or repeated administrations are involved.
Repeatability vs reproducibility
Precision itself splits into two conditions, and the distinction matters because they diagnose different sources of variability:
- Repeatability — agreement between measurements taken by the same operator, on the same instrument, under the same conditions, over a short time interval. This isolates pure random measurement noise.
- Reproducibility — agreement between measurements taken by different operators, on different instruments, in different labs, or on different days. This captures variability introduced by people, equipment, and environment, not just the underlying instrument.
This is the “R&R” in Gauge R&R (Gauge Repeatability and Reproducibility), a standard industrial and lab quality-control study that partitions total measurement variation into repeatability (equipment variation) and reproducibility (operator/condition variation) so each source can be addressed separately. A measurement system can have good repeatability and poor reproducibility — tight results within one operator’s hands, but drift between operators or sites — which is a common finding when a protocol is under-specified. See the reproducibility dictionary entry for how this connects to the broader reproducibility-of-research-findings sense of the term.
Calibration corrects accuracy, not precision
Calibration is the process of comparing an instrument’s readings against a known reference standard and adjusting or documenting the offset. Calibration is how systematic error (bias) gets corrected — it directly targets accuracy.
Calibration does essentially nothing for precision. If an instrument’s readings are noisy and scattered (poor precision), calibrating it against a reference standard will shift the whole scattered cloud of readings toward the true value on average, but it will not make the individual readings agree with each other any better. Precision is a property of the instrument’s internal consistency and the measurement environment; it has to be improved by controlling noise sources (better shielding, more stable operating conditions, a more sensitive or better-designed instrument, tighter operator technique), not by calibration.
Traceability is the property that makes calibration meaningful: an unbroken chain of comparisons linking an instrument’s calibration back to a national or international measurement standard (in the US, maintained by NIST; internationally, coordinated through the BIPM). ISO/IEC 17025 is the general standard governing the competence of calibration and testing laboratories, including the traceability requirement. A calibration certificate without a traceable reference is not verifiable evidence of accuracy.
A different meaning: “accuracy” in diagnostic testing
In diagnostic and classification contexts, “accuracy” is used in a third, distinct sense: the proportion of all classifications (positive and negative) that are correct. This diagnostic-accuracy sense is related to, but not identical with, the metrology sense of accuracy above — and it has a well-known failure mode under class imbalance.
Consider a hypothetical screening test for a condition with a true prevalence of 1% in the population tested, and a test that is 95% accurate in the sense of correctly classifying 95% of both true positives and true negatives (a simplified illustration, not a real study). If everyone in the population is simply classified as “negative” regardless of the test, that trivial rule is already about 99% “accurate” by the correct-classification definition, because 99% of the population truly is negative. A test that actually does some work but still gets 95% of cases right can, depending on how its errors are distributed between false positives and false negatives, still misclassify a large share of the small number of true positives — the group the test exists to find. Overall accuracy averages over both classes and can look excellent while performing poorly on the class that actually matters.
This is exactly why sensitivity and specificity are reported separately from overall accuracy in diagnostic testing: they break performance down by true class (how well positives are caught, how well negatives are correctly cleared) instead of collapsing it into one number that class imbalance can distort. A headline “95% accurate” claim for a rare condition needs sensitivity and specificity reported alongside it to mean anything.
Reporting accuracy and precision correctly
- Significant figures should reflect precision, not exceed it. If an instrument reliably resolves to the nearest 0.1 unit, reporting a result to four decimal places implies a precision the instrument does not have and misleads anyone using the number.
- Report uncertainty, not just a point estimate. A single number without an associated standard deviation, standard error, confidence interval, or stated tolerance gives no way to judge whether it’s precise, let alone accurate.
- Never quote more decimal places than were actually measured. Calculations (averaging, unit conversion) can produce long decimal strings that have no basis in the resolution of the original measurement; round back to what the instrument actually supports before reporting.
- State accuracy and precision separately when both are known. “Accurate to within ±0.5 g, with a repeatability SD of 0.05 g” communicates two genuinely different properties that a single combined figure would obscure.
Frequently asked questions
Is accuracy or precision more important?
Neither is more important in the abstract — they answer different questions, and most applications need both. An instrument that is accurate but imprecise gives an unreliable individual reading; an instrument that is precise but inaccurate gives a confidently wrong one. Which matters more in a given case depends on whether the priority is a single trustworthy reading or a stable, comparable series of readings over time.
Can a measurement be precise but not accurate?
Yes — this is the case of a consistently biased instrument. Repeated readings agree closely with each other (high precision) but are all offset from the true value by roughly the same amount (low accuracy). It is the combination most likely to be mistaken for reliable, because low variability is often used, incorrectly on its own, as evidence of correctness.
What is the difference between repeatability and reproducibility?
Repeatability measures agreement under the same conditions (same operator, instrument, and short timeframe); reproducibility measures agreement across different conditions (different operators, instruments, labs, or days). Both are components of precision, but they diagnose different sources of variability.
Does calibration improve precision?
No. Calibration corrects systematic error (bias) by referencing a known standard, which improves accuracy. It does not reduce the random variability between repeated readings, which is what precision measures — that requires controlling noise sources in the instrument, environment, or procedure.
Why can a “95% accurate” diagnostic test still be unreliable for a rare condition?
Because overall accuracy in classification is the proportion of all cases correctly classified, averaged across both the condition-positive and condition-negative groups. When a condition is rare, the negative group dominates the population, so a test can score very high on overall accuracy while performing poorly on the small positive group that the test is actually meant to detect. Sensitivity and specificity, reported separately, avoid this distortion.







