Sampling & Statistical Power
Before data collection begins, a researcher has to answer two linked questions: how many participants or observations are needed, and how will they be selected. This sub-cluster covers power analysis and sample-size calculation — the a priori estimation of how many observations are needed to detect an effect of a given size at a given significance threshold — alongside the sampling methods (simple random, stratified, cluster, convenience, snowball) that determine whether a sample can support the generalization a study wants to make. Underpowered studies and non-representative sampling are among the most common reasons a study fails peer review or fails to replicate, making this foundational, not optional, groundwork.
Guides
Attrition Bias: When Dropouts Break Your Sample
Attrition bias occurs when study dropouts differ systematically from completers in a way related to the outcome. Learn to distinguish overall from differential attrition, detect it, meet CONSORT reporting requirements, and mitigate it from design through analysis.
Bootstrapping in Statistics: Resampling for Confidence Intervals
Bootstrapping estimates a statistic’s uncertainty by resampling your own data with replacement thousands of times. This guide walks the procedure on a worked example, compares percentile and BCa confidence intervals, gives practical guidance on how many resamples to run, and covers the specific conditions — small samples, dependent data, extreme quantiles — where the method breaks down.
Random Assignment vs Random Sampling: Two Different Randomisations
Random sampling decides who is studied; random assignment decides what happens to them. One builds external validity, the other internal validity — with worked examples for all four combinations.
Selection Bias: How Your Sample Stops Representing Your Population
Selection bias is the umbrella term for systematic distortion in who or what ends up analyzed — sampling bias, self-selection, non-response, attrition, survivorship, and referral bias are its forms. This guide maps where each enters a study and how to detect and mitigate it.
Law of Large Numbers: What Bigger Samples Actually Buy You
What the law of large numbers actually guarantees, the weak vs. strong forms, a worked convergence demonstration, the gambler’s fallacy misreading, and the practical link between sample size and standard error.
Purposive Sampling: Choosing Cases on Purpose
How purposive sampling works, the named variants (maximum variation, homogeneous, typical case, extreme case, critical case, expert), and how to justify the choice in a methods section.
Effect Size: Choosing, Reporting and Interpreting It
A p-value tells you whether a result is unlikely to be chance. It does not tell you whether the result matters. Effect size does — here is how to choose the right measure for your design, avoid Cohen’s benchmark trap, and report it correctly.
Convenience Sampling: When It’s Defensible and When It Isn’t
Convenience sampling is legitimate for pilots, instrument testing, and hard-to-reach populations — but not for population-level claims. Here’s the bias it introduces and how to report it honestly.
The Central Limit Theorem: Why Sample Means Go Normal
The Central Limit Theorem explains why the distribution of sample means approaches normal as sample size increases, regardless of the population’s shape. Covers the formal conditions, a simulation walkthrough, the n≥30 rule of thumb and its limits, and what the theorem does and does not license.
Quota Sampling: Definition, Method, and Examples
How quota sampling works, how it differs from stratified sampling, proportional vs. non-proportional quotas, worked examples, and its advantages and disadvantages.
Cluster Sampling: Definition, Design Effect, and When to Use It
How cluster sampling works, why it trades precision for cost and feasibility, one/two/multistage designs, the design effect (DEFF), correct analysis methods, and when to avoid it.
Sampling Bias: Types, Causes, and How to Detect It
Sampling bias occurs when a sample systematically differs from the population it should represent, distorting results in a direction that more data cannot fix. This guide covers selection, non-response, self-selection, survivorship, healthy-worker, referral, and attrition bias, with real examples and how to detect and report each.
Systematic Sampling: Definition, Method, and When to Use It
Systematic sampling selects a random start between 1 and k, then every k-th unit from an ordered population (k = N/n) — simple to run in the field, but vulnerable to periodicity bias if the list has a hidden cyclical pattern.
Simple Random Sampling: Definition, Steps, and Worked Examples
What simple random sampling actually requires (an equal and independent chance of selection for every population member), the sampling frame problem, how to draw a sample step by step in R, Excel, and Python, worked examples at three scales, and how it compares to stratified, cluster, and non-probability methods.
Power Analysis and Sample Size Calculation: A Complete Guide
How to run a power analysis to calculate sample size: the four interlocking quantities (effect size, alpha, power, N), a priori vs. post-hoc power, choosing an effect size without pilot data, a full worked example, and how to justify N to a reviewer or ethics committee.







