Written and maintained by CASRAI Editorial Board
Last updated
RNA-seq (RNA sequencing) has become the default method for measuring gene expression at genome scale, but most RNA-seq experiments that fail to produce a usable, publishable answer fail before a single library is prepared — they fail at the design stage. Sequencing depth, aligner choice, and normalization method get most of the attention because they are the parts a bioinformatician controls after the fact; replicate structure, batch design, and library strategy are decided by the bench scientist before any sequencing happens, and they are far more consequential to whether the experiment can answer its question at all. This guide works through RNA-seq in the order decisions actually need to be made: experimental design first, then library preparation and RNA quality, then sequencing depth, then the analysis pipeline conceptually, then the research-data-management obligations — storage, deposition, and metadata — that a funded, publishable RNA-seq project now has to satisfy from the outset.
Experimental design: where RNA-seq experiments actually succeed or fail
Biological replicates versus technical replicates
A biological replicate is an independent biological sample — a separate animal, a separate culture dish grown up independently, a separate patient — carried through extraction, library prep, and sequencing separately. A technical replicate is the same RNA extraction (or even the same library) sequenced or processed more than once. Technical replicates tell you about the reproducibility of your assay; they do not tell you anything about whether an expression difference is real across biological variation, and they cannot substitute for biological replicates in a differential expression analysis. This distinction matters because it determines what your statistical test is actually testing: a test run on technical replicates of one biological sample per condition is not a test of the biological effect, no matter how deeply it is sequenced.
The ENCODE Consortium’s RNA-seq guidelines set the field’s reference bar: experiments should be performed with two or more biological replicates unless there is a compelling, stated reason it is impractical, and technical replicates are explicitly not required except to diagnose unusually high biological variability. Two is a floor, not a target. In practice, two replicates gives a differential expression test almost no power to distinguish real biological variance from noise once you account for the fact that RNA-seq count data is over-dispersed (variance exceeds the mean), which is why most methods papers and experienced core facilities recommend three as a common working minimum and describe three as still frequently underpowered for detecting modest fold changes or for tolerating the loss of a sample to QC failure. If the effect size you expect is small, if you need to detect changes across many genes simultaneously (which raises the statistical bar via multiple-testing correction, discussed below), or if you cannot afford to lose a sample and be left with an unbalanced design, plan for four to six biological replicates per condition rather than three.
Batch effects: design them out, don’t try to model them out afterward
A batch effect is any systematic, non-biological source of variation introduced by when or how a sample was processed — different extraction days, different library prep kit lots, different sequencing flow cells or runs, even different technicians. Batch effects are a design problem, not just an analysis problem, because the damage they do is largely a function of whether they are confounded with your biological variable of interest. If every control sample was extracted on Monday and every treated sample on Tuesday, no amount of downstream statistical correction can cleanly separate “treatment effect” from “Tuesday effect” — the two are mathematically inseparable in that dataset. The fix is at the bench: process samples from every condition together, in the same batch, on the same day, with the same reagent lots, where possible; where a batch effect genuinely cannot be avoided (e.g., a study that must run across months), randomize and balance condition assignment across batches so that batch and biology are no longer confounded, and record batch identity explicitly so it can be included as a covariate in the statistical model afterward. Blocking and randomization at the design stage is cheap; trying to statistically model away a fully confounded batch effect after sequencing is, at best, a partial fix and, at worst, not possible.
The real trade-off: more replicates versus more depth
For a fixed sequencing budget, the single most consequential design decision is how to split that budget between sequencing more biological replicates at moderate depth versus fewer replicates sequenced very deeply. The methods literature on this question is reasonably consistent: for standard differential expression analysis, statistical power gained from additional biological replicates substantially outweighs power gained from additional read depth beyond a moderate threshold, because depth increases precision on the expression estimate for a given sample but does nothing to characterize biological (between-sample) variance, and it is biological variance that a differential expression test has to see past. A frequently cited analysis of this trade-off (Liu, Zhang & Zhang, Bioinformatics, 2014, “RNA-seq differential expression studies: more sequence or more replication?”) found that adding replicates identifies substantially more differentially expressed genes than adding equivalent sequencing depth to fewer samples, and that beyond roughly ten reads per gene per sample, additional depth contributes comparatively little to detection power. The practical rule of thumb this supports: when choosing between three deeply-sequenced replicates and five or six moderately-sequenced replicates for the same total cost, the larger replicate number is very often the better use of the budget for a standard differential expression question. This does not hold universally — isoform-level and novel-transcript questions genuinely need more depth per sample, discussed below — but for the most common RNA-seq use case, plan for replication first and treat depth as the variable you trim to fit the budget.
Library preparation choices
Poly-A selection versus ribosomal RNA depletion
Ribosomal RNA makes up the large majority of total RNA in a cell and carries essentially no differential-expression signal, so every RNA-seq library prep method removes it before sequencing, using one of two strategies. Poly-A selection uses oligo-dT beads to capture the poly-adenylated tail present on mature mRNA, which enriches for coding transcripts and depletes rRNA as a side effect; it is simpler, cheaper, and standard for good-quality RNA from fresh or well-preserved samples. It has a real limitation: it will not capture non-polyadenylated RNA species (many long non-coding RNAs, histone mRNAs), and on degraded RNA it produces a pronounced 3′ bias, because fragmentation breaks up the transcript and only fragments still attached to the poly-A tail get captured. Ribosomal RNA depletion instead uses probes or enzymatic digestion (e.g., RNase H-based methods) to remove rRNA directly, regardless of poly-A status, which makes it the correct choice for degraded samples (low RIN or low DV200, see below), FFPE tissue, and any study specifically interested in non-coding or non-polyadenylated transcripts. It costs more and captures more residual rRNA and intronic/pre-mRNA signal than poly-A selection on intact RNA, which is why it is not simply the default choice for every study — it is the correct choice specifically when sample quality or biology requires it.
Stranded versus unstranded libraries
A stranded (strand-specific) library preserves information about which DNA strand a transcript was actually transcribed from; an unstranded library does not. Strandedness matters wherever overlapping genes on opposite strands, antisense transcripts, or accurate quantification of genes with overlapping annotations are relevant — which is most eukaryotic genomes to some degree. Stranded protocols are now close to the default recommendation for new RNA-seq work because the marginal cost over unstranded prep is small and the information, once discarded, cannot be recovered afterward; unstranded remains defensible mainly for simple gene-level counting in well-annotated genomes with little strand-overlap ambiguity, or when reprocessing legacy protocols for direct comparability with older unstranded datasets.
Paired-end versus single-end sequencing
Single-end reads sequence a fragment from one end only; paired-end reads sequence both ends, which improves mapping accuracy (especially across splice junctions and in repetitive regions), is close to a requirement for isoform-level analysis, novel transcript discovery, and fusion/splice-variant detection, and is generally preferred whenever budget allows. Single-end sequencing remains an entirely reasonable, cheaper choice for straightforward gene-level differential expression counting in a well-annotated genome, where the extra mapping precision from paired reads improves quantification only marginally relative to its added cost — another place where the “more replicates over more sequencing” principle above applies: a fixed budget spent on single-end reads across more replicates frequently beats the same budget spent on paired-end reads across fewer.
RNA quality: RIN, DV200, and when quality forces the library-prep decision
RNA integrity is assessed before library prep, most commonly with a Bioanalyzer/TapeStation-style electrophoretic trace scored as a RIN (RNA Integrity Number) from 1 (fully degraded) to 10 (fully intact) — a model score computed across the whole electropherogram, not the 28S:18S ratio it replaced. A RIN of roughly 7 or above is the threshold most commonly cited as suitable for standard poly-A-based RNA-seq; below that, degradation is generally severe enough to distort quantification, and 3′ bias from poly-A selection becomes a real problem. RIN itself becomes unreliable for FFPE (formalin-fixed, paraffin-embedded) tissue, because the fixation process degrades ribosomal RNA in a way that produces artificially low RIN scores that don’t track mRNA usability the same way they do in fresh-frozen tissue. For FFPE and other chemically degraded samples, DV200 — the percentage of RNA fragments longer than 200 nucleotides — is the metric actually built into vendor protocols (including Illumina’s), with roughly 30% DV200 commonly cited as a practical minimum for attempting library construction. The quality metric and the library-prep decision are linked in practice: low RIN or low DV200 samples should generally move to a ribosomal-depletion-based (not poly-A) library prep to avoid compounding degradation with 3′ bias.
Sequencing depth by application
There is no single “right” read depth for RNA-seq — the correct number depends entirely on what question is being asked, and the field’s guidance varies by source and by how comprehensive the answer needs to be. As a general orientation, not a fixed rule:
- Standard gene-level differential expression in a well-annotated genome is the shallowest need — commonly cited ranges run from roughly 10 million up to 30-40 million reads per sample for robust detection of moderate fold changes, with diminishing returns on detection power reported past that range in several published depth-titration studies, reinforcing the replicate-over-depth principle above.
- Isoform-level quantification and alternative splicing analysis need substantially more — guidance in the range of 30-60 million reads and up is common, because distinguishing between transcript isoforms of the same gene requires enough read coverage across exon-exon junctions specifically, not just enough coverage of the gene overall.
- Novel transcript discovery and comprehensive splicing/rare-transcript detection is the deepest tier, with some published analyses citing depths well over 100 million reads (and considerably more for detecting rare differential splicing events specifically) to achieve comprehensive coverage.
Because the numbers reported across the methods literature vary by roughly an order of magnitude depending on genome complexity, desired sensitivity, and exact analytical goal, treat any single figure as a starting point to sanity-check with your sequencing core or bioinformatics collaborator against your specific species, tissue, and question — not as a number to lock in before that conversation.
The analysis pipeline, conceptually
Every standard RNA-seq analysis pipeline, regardless of the specific software chosen, moves through the same conceptual stages:
- Quality control of raw reads (adapter content, base quality, overrepresented sequences, GC content) before anything else happens, so a failed sample or a bad batch is caught before compute time is spent aligning it.
- Alignment or pseudo-alignment — either splice-aware alignment of reads to a reference genome (traditional aligners) or pseudo-alignment/quasi-mapping of reads directly to a transcriptome index (faster, alignment-free methods that estimate transcript-level abundance without producing a full genomic alignment). The choice affects speed and what downstream questions are directly answerable (genome-level splice discovery generally still wants genome alignment), but both are legitimate, widely used approaches for standard expression quantification.
- Quantification — counting (or probabilistically estimating) how many reads/fragments support each gene or transcript.
- Normalization — adjusting raw counts for differences in total sequencing depth and library composition between samples, so that expression levels are comparable across samples rather than reflecting how deeply each sample happened to be sequenced.
- Differential testing — statistical modeling (commonly negative-binomial-based methods, which account for RNA-seq’s characteristic over-dispersion) to identify which genes differ between conditions beyond what would be expected from noise alone.
Why multiple-testing correction is non-negotiable
A typical RNA-seq differential expression analysis tests thousands of genes simultaneously — often 15,000-20,000 or more in a human or mouse study. At a conventional p < 0.05 significance threshold applied independently to every gene, roughly 5% of genes with no real difference will appear statistically significant by chance alone, which at that scale means hundreds of false positives before a single real biological effect has been considered. This is why RNA-seq results are reported as an FDR (false discovery rate), most commonly via Benjamini-Hochberg correction, rather than as raw p-values: it controls the expected proportion of false positives among the genes called significant, rather than the per-gene error rate, which is the correct quantity to control when testing this many hypotheses at once. Any RNA-seq result presented without multiple-testing correction — a gene list filtered on raw p-value alone — should be treated as unreliable regardless of how compelling individual genes on the list look; reviewers and downstream readers should expect adjusted p-values (often labeled “padj” or “q-value”) as the reported statistic, not raw p-values.
The research-data-management layer
RNA-seq generates the kind of large, structured, reusable dataset that funder data policies were written for, and the RDM obligations start well before sequencing — they belong in the data management plan submitted with the grant application, not retrofitted after the sequencer runs.
Data volume and storage planning
Raw sequencing output (FASTQ files) plus downstream alignment and count files add up quickly across a multi-replicate, multi-condition study, and volume scales directly with the depth-times-replicate-count decisions made at the design stage above — another reason those decisions belong in the budget conversation early, alongside sequencing cost itself. A data management plan for an RNA-seq study should state, concretely: an estimate of total raw and processed data volume across the full sample set, where raw and processed data will be stored during active analysis, backup and retention arrangements, and who is responsible for storage costs once active grant funding ends — storage is not a one-time cost, and reviewers increasingly expect this stated explicitly rather than assumed.
Deposition: GEO, SRA, and ArrayExpress
Public deposition of RNA-seq data is now close to a universal journal and funder requirement, not an optional courtesy. In the US, this typically means the Gene Expression Omnibus (GEO) for processed expression data and sample-level annotation, which is itself built on top of raw-read deposition to the Sequence Read Archive (SRA) — a GEO RNA-seq submission is structured as linked BioProject, BioSample, and SRA experiment/run records, all now handled through NCBI’s single Submission Portal. The European equivalent is ArrayExpress at EMBL-EBI. Compliance with these repositories’ data standards is governed by MINSEQE (Minimum Information About a Sequencing Experiment), the sequencing-era counterpart to the older microarray-focused MIAME standard, both originating with what is now the FGED Society. MINSEQE compliance is about content, not submission format: it requires raw data, processed data, sample annotation, a description of the experimental design, feature/platform annotation, and the protocols used — all six need to be captured during the experiment, not reconstructed from memory at submission time.
Embargo options
GEO and SRA both support holding a deposited dataset as private (embargoed) until a specified release date or until the associated manuscript is published, rather than requiring immediate public release at deposition — this satisfies the common funder/journal requirement to have data deposited (and an accession number available to cite in the manuscript) prior to publication, while still protecting the dataset from public access until the paper itself is out.
When human RNA-seq needs controlled access instead of open deposition
Human RNA-seq data is not automatically safe for open, unrestricted deposition the way model-organism data usually is: expressed transcripts can carry enough incidental genetic variation to make re-identification a realistic risk, particularly at higher sequencing depth. Where informed consent or the nature of the data requires it, the appropriate route is controlled access via dbGaP (NIH’s Database of Genotypes and Phenotypes) rather than the public SRA/GEO — dbGaP applies a Data Access Committee review process for each downstream requester rather than open access, and functional-genomics-only results without genotype-linkable content may still route to GEO even when raw reads require dbGaP; check this distinction with your repository and your IRB before deposition, not after. Human genomic data governed by the NIH Genomic Data Sharing Policy needs this determination made at the consent and data management plan stage, since it affects what a study can legally promise a participant about how their data will be shared.
Core facility recharge and budgeting on a grant
Most institutions run RNA-seq through a shared sequencing core facility rather than in individual labs, on a recharge (cost-recovery) basis. Budgeting an RNA-seq study on a grant should account for the full chain, not just the per-sample sequencing cost quoted by the core: RNA extraction and QC (RIN/DV200 assessment), library preparation (which varies significantly by strategy — poly-A versus rRNA depletion, stranded versus unstranded change the per-sample reagent cost), the sequencing run itself (priced by depth and read length, and often cheaper per sample when multiple studies’ libraries are pooled and multiplexed onto a shared flow cell/lane), bioinformatics analysis support if the core or a separate facility provides it, and data storage for the life of the grant plus any required post-award retention period. Request an itemized quote from the core facility early, before finalizing replicate number and depth in the grant budget — core facilities can generally advise on realistic depth-per-dollar figures for your organism and application far more precisely than a general guideline can, and getting that quote before submission avoids a mismatch between what the grant funds and what the experiment actually needs.
Frequently asked questions
Is 3 biological replicates enough for RNA-seq?
Three meets the commonly cited working minimum and exceeds ENCODE’s stated floor of two, but three is frequently underpowered for detecting modest fold changes, tolerates no sample loss without becoming unbalanced, and gives limited power once multiple-testing correction is applied across thousands of genes. Where budget allows, four to six replicates per condition is a more defensible target for a standard differential expression design.
Should I prioritize more replicates or more sequencing depth?
For standard differential expression analysis on a fixed budget, published depth-titration analyses generally favor more biological replicates over deeper sequencing of fewer samples, because depth improves precision on a single sample’s estimate while replicates are what let a statistical test distinguish real biological variance from noise. Isoform-level and novel-transcript questions are the exception and genuinely need more depth per sample.
Do I need poly-A selection or ribosomal RNA depletion?
Poly-A selection is the standard, cheaper default for good-quality RNA (RIN roughly 7+) from fresh or well-preserved samples. Ribosomal RNA depletion is the correct choice for degraded samples, FFPE tissue, and studies specifically interested in non-polyadenylated or non-coding RNA.
What RIN score do I need for RNA-seq?
A RIN of roughly 7 or higher is the threshold most commonly cited as suitable for standard poly-A-based RNA-seq. For FFPE and other chemically degraded samples, RIN itself becomes unreliable and DV200 (percentage of fragments over 200 nucleotides, with roughly 30% cited as a practical minimum) is the metric actually used to assess whether library construction is feasible.
Do I have to deposit RNA-seq data publicly?
Most funders and journals now require deposition of RNA-seq data to a recognized public repository (GEO/SRA in the US, ArrayExpress in Europe) as a condition of publication, with an embargo option available to delay public release until the associated paper is published. Human data that could be re-identifying, or that is governed by the NIH Genomic Data Sharing Policy, may need to route through controlled-access dbGaP instead of open deposition — confirm this with your IRB and repository before deposition.
What’s the difference between GEO and SRA?
SRA holds raw sequencing reads; GEO holds processed expression data and sample-level experimental annotation, and a GEO RNA-seq submission is built on top of linked SRA raw-read records (BioProject, BioSample, SRA experiment/run) submitted through NCBI’s shared Submission Portal. See the GEO guide for the full submission mechanics.
How is RNA-seq different from RT-qPCR?
RNA-seq measures expression genome-wide without needing to know in advance which genes to look at; RT-qPCR measures a small, predefined panel of targets with higher per-gene precision at much lower cost, and is the standard method for validating a shortlist of RNA-seq hits. See the RT-qPCR guide for validation-specific design considerations, including reference-gene selection.
Further reading: Cluster randomized Trials — What the intracluster correlation coefficient (ICC) measures in a cluster randomised trial, how it produces the design effect that inflates required sample size, why the number of clusters matters more than total partici.








