Skip to main content
v2026.11,610 entries · CC-BY 4.0

Net Reclassification Improvement (NRI): The Calculation, and the Case Against It

The NRI compares two risk prediction models by counting who moved in the right direction. This guide gives the arithmetic for both the category-based and category-free versions with a worked reclassification table, then sets out the published methodological case against the statistic — including simulations in which a marker with no predictive information at all produced a positive, statistically significant NRI.

Ask about Net Reclassification Improvement (NRI): The Calculation, and the Case Against It

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

The net reclassification improvement (NRI) was introduced by Pencina and colleagues in Statistics in Medicine in 2008 to answer a real complaint: the area under the ROC curve barely moves when a genuinely useful biomarker is added to an already-good risk model. The NRI spread fast, and by the mid-2010s it was near-standard in cardiovascular and oncology biomarker papers.

It also attracted an unusually specific and unusually damaging methodological literature. A 2014 paper in the Journal of the National Cancer Institute ended with the sentence: “Use of NRI P values in scientific reporting should be halted.” That is not a hedged caveat, and it has not been retracted.

This guide does both halves. It gives the arithmetic — category-based and category-free — with a worked reclassification table you can follow number by number, and then it gives the case against, with the primary citations attached to each specific objection. The point is not to compute an NRI. The point is to let you read an NRI claim in someone else’s paper and decide how much of it to believe.

What the NRI is actually measuring

The NRI compares two risk prediction models — a baseline model and the same model plus a new marker — by asking how many people moved in the right direction when the marker was added.

“Right direction” is defined separately for people who go on to have the event and people who do not:

  • For an event (a case), a predicted risk that goes up is an improvement.
  • For a non-event (a control), a predicted risk that goes down is an improvement.

The NRI is the sum of two net-movement proportions, one computed within events and one within non-events:

NRI = [(events moving up − events moving down) / nevents] + [(non-events moving down − non-events moving up) / nnon-events]

Each bracketed component ranges from −1 to +1, so the NRI itself ranges from −2 to +2. This is the first thing that trips people up: it is not a proportion of the study population, and it cannot be read as one. More on that below.

Two variants exist, and they are different statistics that happen to share a name:

  • Category-based NRI (Pencina et al., 2008) — movement is only counted when a person crosses a pre-defined risk-category boundary, e.g. the 5% / 20% ten-year cardiovascular risk thresholds.
  • Category-free, or continuous, NRI (Pencina, D’Agostino and Steyerberg, 2011) — any change in predicted risk counts, however small, in whichever direction it goes.

Pencina and colleagues stated plainly in the 2011 paper that “NRIs cannot be compared across studies unless they are defined in the same manner.” An NRI reported without saying which version it is, and with what cut-offs, is uninterpretable on its face.

Category-based NRI: a worked example

Illustrative composite. The cohort and counts below are constructed to make the arithmetic followable; they are not drawn from any real study and should not be cited as data.

A cohort of 1,000 people, followed for ten years. 100 have the event (10% event rate); 900 do not. Risk is categorised as <5%, 5–20%, >20% — the classic three-category cardiovascular split. Every person has a predicted risk under the baseline model and under the baseline-plus-marker model, and we count who crossed a boundary.

Step 1 — the events (n = 100)

Movement Count
Moved up a category (correct direction) 32
Moved down a category (wrong direction) 12
Stayed in the same category 56

NRIevents = (32 − 12) / 100 = +0.20

Step 2 — the non-events (n = 900)

Movement Count
Moved down a category (correct direction) 54
Moved up a category (wrong direction) 99
Stayed in the same category 747

NRInon-events = (54 − 99) / 900 = −45 / 900 = −0.05

Step 3 — add them

NRI = 0.20 + (−0.05) = +0.15

What that number does and does not say

Read the components, not the total. The headline “NRI = 0.15” is entirely carried by the events. Among the 900 non-events — 90% of the cohort — the new marker made things net worse: 99 people were pushed into a higher risk category they did not belong in, against 54 correctly moved down. In a screening context, those 99 are additional statins, additional imaging, additional follow-up, for people who were never going to have the event.

The two components are weighted equally in the sum even though they are computed over populations that differ by a factor of nine. That is a deliberate design choice, not an error — but it means the NRI total silently trades a large number of false alarms for a small number of correct upgrades, and the total alone will not tell you it did. Kerr and colleagues recommended in Epidemiology in 2014 that the components always be reported separately for exactly this reason.

The event component and the non-event component are also, when there are only two risk categories, the same quantities as the change in the true-positive rate and false-positive rate. Kerr and colleagues argued the existing, descriptive terms are the more useful ones and should simply be retained.

Category-free (continuous) NRI

The category-free NRI removes the boundaries. Every person whose predicted risk moved in the right direction counts, no matter how far they moved.

Using the same illustrative cohort: suppose 66 of the 100 events had their predicted risk go up under the new model and 34 go down, while 430 of the 900 non-events went down and 470 went up.

  • Event component: (66 − 34) / 100 = +0.32
  • Non-event component: (430 − 470) / 900 = −0.044
  • Continuous NRI = +0.276

Note what has happened. The number got substantially bigger — 0.28 versus 0.15 — on the same data, for the same marker, because a person whose risk moved from 8.0% to 8.1% now counts exactly as much as a person who moved from 8% to 40%. That insensitivity to magnitude is the whole mechanism. It is why the category-free NRI is systematically larger than the category-based version and why the two must never be compared to each other.

The category-free NRI was proposed partly to make NRIs comparable across studies by removing arbitrary cut-off choices. Kerr and colleagues concluded the opposite: that it “can mislead investigators by overstating the incremental value of a biomarker, even in independent validation data” — the italicised clause being the part that matters, because independent validation is usually the defence offered when this criticism is raised.

The IDI, its close relative

The integrated discrimination improvement (IDI) came from the same 2008 Pencina paper and travels with the NRI in most papers. Rather than counting movements, it averages them:

IDI = (mean predicted risk change among events) − (mean predicted risk change among non-events)

It is equivalent to the improvement in the discrimination slope — the gap between the mean predicted risk in events and in non-events. It shares the NRI’s problems, and in the case of its variance estimate the evidence against it is older: Kerr, McClelland, Brown and Lumley showed in the American Journal of Epidemiology in 2011 that the published standard-error method “tends to underestimate the error,” and that the proposed z test “is not valid, because the null distribution of the test statistic is not standard normal, even in large samples.”

The methodological case against the NRI

The objections below are not stylistic preferences. Each is a published, simulation-backed or theorem-backed result, and each is cited to the paper that established it.

1. A marker with no predictive information can produce a positive NRI

This is the most serious charge and it holds up. Pepe, Fan, Feng, Gerds and Hilden, writing in Statistics in Biosciences in 2015, demonstrated “the alarming result that the NRI statistic calculated on a large test dataset using risk models derived from a training set is likely to be positive even when the new marker has no predictive information.” They also gave a theoretical example in which an incorrect risk function containing an uninformative marker is proven to yield a positive NRI.

Note the setup carefully: the NRI is computed on a large, independent test dataset, using models fitted on separate training data. That is the design most readers would accept as a clean validation. It is not sufficient here. The authors’ explanation is that a large NRI can arise purely from the risk models fitting poorly, which is a property the NRI does not penalise. They contrast this with measures derived from the ROC curve, the net benefit function, and the Brier score, which “cannot be large due to poorly fitting risk functions.”

2. Poor calibration makes a model look better, not worse

Hilden and Gerds, in Statistics in Medicine in 2014, showed that if the IDI and NRI are used to measure gain in prediction performance, “poorly calibrated models may appear advantageous,” and — the sharper finding — that in simulation “even the model that actually generates the data (and hence is the best possible model) can be improved on without adding measured information.”

The technical property being described here is propriety. A scoring rule is proper if its expected value is optimised by reporting your honest best estimate of the probability; the Brier score and the logarithmic score are proper, which is why a well-calibrated model cannot be beaten on them by distorting its own outputs. The NRI is not proper. Hilden and Gerds do not use the word “proper” in that abstract, so we are naming the property rather than quoting them on it — but a demonstration that the data-generating model can be improved upon without new information is precisely a demonstration of impropriety. Their own summary of the contrast is that AUC and the Brier score have “the characteristic that prognostic performance cannot be accidentally or deliberately inflated.” The NRI does not have that characteristic.

The practical consequence: a model whose calibration you have not checked can post a good NRI because its calibration is bad. Calibration reporting is therefore not an optional extra alongside an NRI claim — it is load-bearing for whether the NRI means anything at all. Gerds and Hilden went further in a follow-up note whose title is the argument: calibration of models is not sufficient to justify NRI.

3. The published variance formula produces invalid p-values

Pepe, Janes and Li ran the direct experiment and published it in JNCI in 2014. They built a population of 10,000 individuals with a 10.2% event rate, in which four biomarkers had no predictive ability at all, then repeatedly drew training samples (n = 420) and test samples (n = 420 or 840) and asked how often the NRI came back positive and statistically significant.

Statistic used False-positive rate on markers with zero predictive ability
NRI, evaluated on the training data 63.0%
NRI, evaluated on independent test data 18.8% – 34.4%
Change in AUC Rare
Likelihood ratio statistic ~5.0% (i.e. as expected)

A nominal 5% test that fires 63% of the time on pure noise is not a slightly conservative test; it is not a test. Even the independent-test-data figures are three to seven times the nominal rate. Their conclusion: “Conclusions about biomarker performance that are based primarily on a statistically significant NRI statistic should be treated with skepticism. Use of NRI P values in scientific reporting should be halted.”

Kerr and colleagues reached the compatible practical recommendation: if you are going to report an NRI at all, confidence intervals “should be calculated using bootstrap methods rather than published variance formulas.”

4. Testing whether the NRI exceeds zero is redundant anyway

Pepe, Kerr, Longton and Wang proved in Statistics in Medicine in 2013 that the null hypothesis of no improvement in prediction performance is equivalent to the simple null hypothesis that the new marker Y is not a risk factor once you control for the baseline predictors X — formally, H0: P(D=1 | X,Y) = P(D=1 | X).

If that is true, then testing for improvement in prediction performance is redundant once the marker has been shown to be a risk factor in the model. The well-established tests for regression coefficients already answer the question, and they answer it with correct type I error. Their recommendation is that hypothesis testing be confined to evaluating the marker as a risk factor, and that analyses of prediction performance “focus on estimation rather than on testing for no improvement.”

The same paper contains a finding that cuts the other way and is worth knowing, because it undermines the original motivation for inventing the NRI: standard AUC-comparison procedures that do not adjust for variability in the estimated regression coefficients are “extremely conservative.” The widespread belief that the AUC is insensitive to real improvements may therefore be an artefact of invalid inference procedures rather than a defect of the AUC itself.

5. With three or more categories, clinically unequal moves count equally

In the three-category worked example above, a non-event moving from >20% down to 5–20% and a non-event moving from >20% all the way down to <5% both contribute exactly 1 to the same count. So do the corresponding upward moves. Kerr and colleagues recommended against category-based NRIs with three or more categories on precisely this ground: they “do not adequately account for clinically important differences in shifts among risk categories.”

With only two categories the objection does not apply, because there is only one possible move — but then, as noted, the components are just the changes in true- and false-positive rates, and there is little reason to rename them.

6. “NRI = 0.15” does not mean 15% of patients were reclassified

This misreading is common enough that Leening, Vedder, Witteman, Pencina and Steyerberg listed it explicitly in their Annals of Internal Medicine review. They examined 67 publications in high-impact general clinical journals that used the NRI and found “incomplete reporting of NRI methods, incorrect calculation, and common misinterpretations.” Their systematic recommendation includes, verbatim, “do not interpret the overall NRI as a percentage of the study population reclassified.”

It cannot be that percentage: it is a sum of two proportions computed over two different denominators, and it can exceed 1. In the worked example, +0.15 arose from 44 events and 153 non-events actually changing category — 197 people, or 19.7% of the cohort, of whom a majority moved the wrong way. Note also that this review is co-authored by Pencina, who introduced the statistic; the corrective literature is not entirely external to it.

What the critics recommend instead

The critique is not purely negative. Across these papers the recommended alternatives are consistent:

  • Report changes in true-positive and false-positive rates at clinically meaningful thresholds, in their ordinary names. See sensitivity, specificity and ROC curves for how those are constructed and what the 2×2 table gives you.
  • Net benefit and decision curve analysis. Kerr and colleagues named improvement in net benefit “the preferred single-number summary of the prediction increment,” because it weights false positives against true positives using an explicit, stated threshold probability instead of an implicit one.
  • The Brier score, which is a proper scoring rule and therefore cannot be inflated by miscalibration.
  • A likelihood ratio test on the marker’s coefficient in the model — the test that behaved correctly at 5% in the JNCI simulation, and the one Pepe et al. (2013) showed answers the same question.
  • Calibration assessment, reported explicitly, since the NRI critique turns substantially on calibration being unexamined.

Note that “net benefit” addresses a question the NRI never asked: whether acting on the model helps, at a stated threshold. That is closer to the distinction between statistical and clinical significance, and it is the reason decision-analytic measures have largely displaced reclassification indices in careful prediction-model papers. For the downstream question of what a threshold change costs and buys in patients, absolute risk reduction and number needed to treat are the relevant currency; for the operational version of choosing and calibrating a cut-off in practice, see early warning score implementation.

Appraising an NRI claim in a paper: a checklist

Run these in order. The first four are usually answerable from the methods section in under a minute, and they resolve most cases.

  1. Which NRI is it? Category-based or category-free. If the paper does not say, stop — the number is uninterpretable and cannot be compared with any other paper’s.
  2. If category-based, what are the cut-offs, and were they pre-specified? Cut-offs chosen after seeing the data are a researcher degree of freedom that directly moves the statistic.
  3. Are both components reported separately? Events and non-events. If only the total is given, you cannot tell whether the gain came from correctly upgrading cases or from a flood of false alarms among controls. This is Kerr et al.’s primary reporting recommendation and it is frequently ignored.
  4. Was the NRI computed on the same data the model was fitted on? If yes, the JNCI simulation’s 63% false-positive rate is the relevant reference point, and the result carries essentially no evidential weight.
  5. Is calibration reported for both models? If not, the Hilden and Gerds objection is live and unaddressed: the improvement may be an artefact of the baseline model fitting badly.
  6. How was the confidence interval computed? Bootstrap, or the published asymptotic variance formula? The latter is known to be wrong.
  7. Is a p-value attached to the NRI, and is it doing work in the conclusions? If the abstract’s claim rests on “NRI = 0.14, p = 0.03,” you are looking at exactly the practice the JNCI paper said should be halted.
  8. Is the marker a significant predictor in the model itself? If it is, the NRI adds no inferential information (Pepe et al., 2013). If it is not, an apparently significant NRI is a contradiction that needs explaining, not celebrating.
  9. Is any decision-analytic measure reported alongside — net benefit, a decision curve, the Brier score? Their presence is a strong signal the authors know this literature. Their absence is not fatal, but it shifts the burden.
  10. Does the paper describe the NRI as “x% of patients reclassified”? If so, the authors have misunderstood their own statistic, and that should colour how you read the rest of the analysis.

A paper that passes 1–5 and reports a modest, component-wise, bootstrap-interval NRI alongside a decision curve is doing something defensible. A paper whose headline finding is a significant category-free NRI p-value, computed on the development data, with no calibration reported, is making a claim the primary methodological literature has specifically said cannot be supported.

What the reporting guidelines require

Prediction-model studies fall under TRIPOD (Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis) and, for models with a machine-learning component, TRIPOD+AI. TRIPOD requires reporting of both discrimination and calibration, and of the model’s performance measures with confidence intervals — which is why the calibration objection above maps onto an existing reporting requirement rather than an extra demand. Diagnostic accuracy studies fall under STARD. See STARD and TRIPOD+AI for what each checklist covers and which one applies to your study design.

Neither checklist mandates or forbids the NRI. That is the point worth internalising: a paper can be fully TRIPOD-compliant and still report an NRI in a way the statistical literature says is uninformative. Compliance with a reporting guideline is a transparency guarantee, not a validity guarantee.

Survival data and case-control data

The 2011 Pencina, D’Agostino and Steyerberg paper set out a prospective formulation of the NRI that extends to survival and competing-risk data, and showed how the NRI can be applied to case-control data. In the survival setting the “events” and “non-events” are defined at a fixed horizon and censoring has to be handled explicitly, typically with Kaplan-Meier or inverse-probability weighting — see how to read a Kaplan-Meier curve for what censoring does to the risk sets underneath.

Everything in the case against applies to these extensions as well; the estimation problem is strictly harder, not easier, once censoring weights are involved. The event rate also changes the NRI directly, and Pencina and colleagues flagged that differing event rates across samples, event definitions and follow-up durations make cross-study NRI comparisons unsafe even before the objections above are considered.

Frequently asked questions

Is the NRI valid?

As a descriptive summary of how a reclassification table moved, computed on independent data, reported component-wise, alongside calibration, with a bootstrap interval and no p-value — it is defensible, and Kerr et al. explicitly allow for investigators who find it useful. As an inferential statistic used to establish that a biomarker adds predictive value, the published evidence is that it is not valid: it can be positive for markers with no predictive information, and its p-values have false-positive rates several times nominal.

What is a good NRI value?

There is no threshold, and treating any particular value as a benchmark is a category error given the calibration dependence. Category-free NRIs are systematically larger than category-based ones on the same data, so a “large” number may only mean the continuous version was used. Judge the components and the study design, not the magnitude.

What is the difference between NRI and IDI?

The NRI counts how many people moved in the right direction; the IDI averages how far predicted risks moved, and is equivalent to the change in the discrimination slope. They come from the same 2008 paper, they are usually reported together, and the same criticisms — miscalibration inflating them, invalid variance estimates — apply to both.

Why not just use the change in AUC?

The original argument for the NRI was that the AUC is insensitive to real improvements. Pepe et al. (2013) found the insensitivity may largely be a property of the testing procedure rather than the measure, since standard AUC comparisons that ignore variability in estimated coefficients are extremely conservative. The AUC also has the property the NRI lacks: it cannot be inflated by a poorly fitting risk model.

Can I compare an NRI from one study to another?

Only if both used the identical definition, the identical cut-offs, and comparable event rates and follow-up. Pencina and colleagues said so themselves in 2011. In practice this condition is almost never met, which is one reason NRIs do not aggregate into meta-analysis usefully.

Should I report an NRI at all?

If a journal or reviewer requires it, report it as a description — both components, computed on validation data, with a bootstrap interval, with calibration shown, and with no p-value — and put a decision-analytic measure such as net benefit next to it as the measure your conclusions actually rest on. Do not let the NRI carry the inferential weight of the paper.

Primary sources

Each claim above is attributable to one of these. Where this guide characterises rather than quotes a paper — notably the word “proper” in the scoring-rule discussion — that is flagged in the text.

A note on where the critique actually lives: it is often described as a Statistics in Medicine and American Journal of Epidemiology literature. That is only partly right. The Statistics in Medicine papers (Pepe 2013, Hilden and Gerds 2014) and the AJE paper (Kerr 2011, on the IDI) are real and load-bearing, but the two most damaging NRI-specific results are in Statistics in Biosciences (2015) and the Journal of the National Cancer Institute (2014), and the most complete practical review is in Epidemiology (2014). If you are chasing these down, search by author rather than by journal.

For related material on model evaluation and measurement, see logistic regression interpretation and diagnostics, criterion validity, and the research methods hub.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →