Written and maintained by CASRAI Editorial Board
Last updated
Most qPCR normalization strategies start and end with a single gene — usually GAPDH, ACTB or 18S rRNA — chosen because it is the one the lab has always used, not because it was shown to be stable in the samples actually being compared. That shortcut works until the reference gene itself moves with the treatment, and at that point every fold-change in the dataset is wrong in a direction nobody can see from the data alone: a target that looks unchanged may be genuinely down, and a target that looks induced may be flat. Reference-gene selection is an experiment, not a default, and it needs the same validation step as any other measurement the assay depends on.
Why an unvalidated reference gene is a silent failure mode
A gene is not a housekeeping gene because it is labeled one. GAPDH is a glycolytic enzyme and its transcription responds to hypoxia and metabolic state; ACTB expression shifts with proliferation rate and with treatments that affect the cytoskeleton; 18S rRNA is so abundant relative to mRNA targets that small extraction or reverse-transcription differences distort it disproportionately, and it is frequently unsuitable for normalizing against oligo-dT-primed cDNA in the first place. None of that makes these genes unusable — it makes them candidates that have to earn their place in a specific experiment, the same way any other reference measurement would.
The failure is silent because a normalization artifact looks exactly like a real result. If the reference gene shifts 1.5-fold between two conditions and the analysis assumes it is flat, every target in the dataset inherits that 1.5-fold error, consistently, in the same direction. There is no internal signal in a delta-delta Ct table that flags this — it only shows up when someone re-runs the analysis with a validated reference and the fold-changes move. See how RT-qPCR works and building a qPCR standard curve for the upstream steps this normalization step sits on top of.
What MIQE requires, and what it doesn’t settle for you
The MIQE guidelines (Bustin et al., Clinical Chemistry, 2009) list reference gene validation as essential reporting, not optional context — MIQE 2009 explicitly required disclosing the number of reference genes used and, where relevant, how they were selected and validated. The 2025 revision, MIQE 2.0 (Bustin, Ruijter, van den Hoff et al., Clinical Chemistry, 2025), replaced the old essential/desirable split with a single unified Yes/No checklist spanning reagent preparation through data analysis, and reference gene justification remains part of it. See the full checklist in MIQE guidelines: the qPCR reporting checklist.
What MIQE does not do is hand you a validated gene list. It tells you validation has to happen and has to be reported; it does not pick genes for you, because stability is sample-set-specific — a panel validated in one cell line, tissue or treatment arm is not automatically valid in a different one. The rest of this guide is the workflow that fills that gap.
Building the candidate panel
Start from 8–12 candidate genes, not the two or three a bench protocol usually names by default. A wider starting panel matters because stability is empirical, not assumed: a gene that looks stable in the literature can behave differently in your specific cell type and treatment. Draw candidates from different functional classes so a single confound (e.g., anything downstream of glycolysis) cannot compromise the whole panel at once. Commonly used candidates include GAPDH, ACTB, B2M, HPRT1, TBP, YWHAZ, RPL13A, PPIA (cyclophilin A), UBC, SDHA, GUSB and 18S rRNA — treat this as a starting list to screen, not a shortlist to trust.
Run every candidate across a sample set that actually represents the comparison the study will make — every treatment arm, timepoint, genotype or tissue you plan to compare, not just untreated controls. A panel validated only in control samples tells you nothing about whether those genes stay stable once the treatment is applied, which is precisely the condition the normalization has to hold up under.
geNorm: pairwise variation and the minimum-gene rule
geNorm (Vandesompele et al., Genome Biology, 2002) ranks candidates by an M-value: the average pairwise variation of each gene against every other candidate in the panel, calculated on log-transformed relative quantities. Genes with the lowest M-values are the most stable; the algorithm works by stepwise exclusion, dropping the least stable gene and recalculating until only the best-performing candidates remain.
geNorm also answers a question raw ranking doesn’t: how many genes are enough? It calculates a pairwise variation value, Vn/n+1, comparing the normalization factor built from the n most stable genes against the factor built from n+1. The original paper proposes 0.15 as the cutoff below which adding another gene is not worth it. In practice this usually settles on two to four genes, not one — which is itself the core argument against single-gene normalization: geNorm’s own algorithm treats one gene as provisional by construction, never as a final answer.
The output you actually use downstream is not a single “winner” gene but a normalization factor: the geometric mean (not arithmetic mean) of the relative quantities of the genes the V-value analysis says to keep. Geometric averaging is the deliberate choice in the original method — it controls for outliers and for differences in absolute expression level between genes in a way an arithmetic mean does not.
NormFinder: a different way to model the same question
NormFinder (Andersen, Jensen & Ørntoft, Cancer Research, 2004) approaches stability differently: instead of pairwise comparison, it fits a model-based estimate of expression variation and explicitly separates within-group variation from between-group variation across the sample groups you define (e.g., treated vs. control, or multiple tissue types). That distinction matters for exactly the failure mode this guide opened with — a candidate gene can look stable overall while actually shifting systematically between the groups being compared, which is the specific pattern that corrupts a fold-change result. NormFinder is built to surface that pattern rather than average over it, and it can also propose the best two-gene combination rather than a single best gene.
geNorm and NormFinder do not always agree on the top-ranked gene, because they are optimizing for different things — pairwise co-variation versus intra-/inter-group variance. Disagreement between the two is informative, not a bug: it usually means at least one candidate is inconsistent enough that its exact rank is unstable, which is itself a reason to keep more than one reference gene rather than commit to whichever tool’s single top pick you ran first.
Other validation tools worth knowing
Two others come up often enough to name. BestKeeper (Pfaffl et al.) works directly from raw Cq values rather than converted quantities, ranking candidates by standard deviation and by their pairwise correlation with each other and with a computed “BestKeeper index.” RefFinder is a web-based aggregator that runs geNorm, NormFinder, BestKeeper and the comparative delta-Ct method together and produces a single consensus ranking by combining their individual rank outputs — useful specifically when the individual tools disagree and you want one defensible ranking to report, rather than picking whichever method’s answer you liked best.
Calculating and using the normalization factor
Once the validation step has settled on a set of stable genes (commonly two to four, per the geNorm V-value cutoff above), the normalization factor applied to every target gene’s Cq is the geometric mean of the selected reference genes’ relative quantities, not any single gene’s Cq and not an arithmetic average. This is the step that actually gets used in the delta-delta Ct or Pfaffl calculation downstream — the reference-gene validation work described above exists specifically to justify what goes into this one number.
Common mistakes
- Validating on control samples only. A panel that looks stable in untreated cells tells you nothing about stability once the actual experimental perturbation is applied — validate across every group the study will compare, not a convenient subset of them.
- Starting from too narrow a candidate list. Screening only GAPDH, ACTB and 18S rRNA and picking “the best of three” is not the same experiment as screening 8–12 candidates from different functional classes — a narrow panel can only ever tell you which of a few bad options is least bad.
- Treating a published validation as portable. A reference-gene panel validated in someone else’s cell line, tissue or species is a starting candidate list for your own validation, not a substitute for running it — stability is sample-set-specific by definition.
- Reusing a validated panel indefinitely. A panel validated for one project’s conditions does not automatically transfer to a new treatment, timepoint or cell passage range added later without re-checking stability under the new conditions.
- Reporting a normalization method without reporting the validation. MIQE-consistent reporting states which genes were tested, which tool(s) were used, and which were selected — not just “normalized to GAPDH,” which gives a reviewer no way to evaluate whether that choice was justified.
Frequently asked questions
Can I just use GAPDH as my qPCR reference gene?
Only if you have validated it as stable across the specific samples and conditions in your study — GAPDH is a glycolytic enzyme whose transcription is known to respond to hypoxia and metabolic state, so “everyone uses it” is not evidence it is stable in your system. Screen it alongside other candidates rather than assuming it in advance.
How many reference genes does a qPCR experiment need?
geNorm’s pairwise variation analysis (the Vn/n+1 value, with 0.15 as the original proposed cutoff) is the standard way to answer this for your own dataset; in practice it usually lands on two to four genes rather than one. There is no universal fixed number — it depends on how stable the specific candidates turn out to be in your samples.
What is the difference between geNorm and NormFinder?
geNorm ranks candidates by average pairwise variation against every other candidate (an M-value) and uses stepwise exclusion. NormFinder instead fits a model-based variance estimate that explicitly separates within-group from between-group variation across your defined sample groups. They often agree on which genes are stable but can rank the single “best” gene differently, because they are measuring different things.
Do geNorm and NormFinder need special software?
Both were originally released as standalone tools (geNorm, NormFinder) and are also implemented inside several commercial qPCR analysis packages and R packages; RefFinder is a web tool that runs both alongside BestKeeper and the comparative delta-Ct method and returns a combined ranking.
Does MIQE require a specific reference-gene validation method?
No. MIQE requires that reference gene selection be justified and reported — the number of genes used and, where applicable, how they were validated — not any one specific tool. geNorm, NormFinder, BestKeeper and RefFinder are all accepted ways to produce and document that justification. See MIQE guidelines: the qPCR reporting checklist for the full reporting requirements.








