Computing & Data
Data Science & Statistics
Statistical methodology and applied data science across domains. Sits closest to the research-methods cluster: study design, inference, and analysis practice, plus the tooling and reproducible-workflow infrastructure that supports them.
23 pages across 4 topic areas
Research Methods & Statistics
- Test-Retest Reliability: Choosing the Retest Interval, and Checking for Drift
The retest interval is the design decision a test-retest coefficient cannot recover from. What the interval trades off, what COSMIN actually asks abou…
- Cook’s Distance and Influential Observations: Competing Thresholds, Leverage vs DFBETAS, and What to Do Next
The three Cook’s distance thresholds in circulation disagree with each other and none is what Cook proposed. A worked R demonstration separating outli…
- The Kolmogorov-Smirnov Test: One-Sample vs Two-Sample, and the Estimated-Parameter Trap
The one-sample Kolmogorov-Smirnov test and the two-sample test answer different questions and have different validity conditions. The one-sample versi…
- Non-Response Bias: How to Measure It and What to Report
Response rate is a weak proxy for non-response bias. Three measurement procedures — wave analysis, a non-respondent follow-up sub-sample, and benchmar…
- Difference-in-Differences: Evidencing Parallel Trends and the Staggered-Adoption Problem
Parallel trends is a counterfactual assumption, not something a pre-trend plot can confirm. This guide sets out what each diagnostic actually establis…
- Endogeneity: The Three Sources, and the Remedy That Matches Each
Endogeneity is three different problems – omitted-variable bias, simultaneity and measurement error – and the remedy that fixes one does nothing for t…
- The Wilcoxon Signed-Rank Test: Assumptions, Exact vs. Normal Approximation, and How to Report It
A decision guide to the Wilcoxon signed-rank test: what its null actually asserts (symmetry of the differences, not equal medians), why it is the pair…
- Inter-Rater Reliability: Choosing the Right Coefficient
A selection table from data type and rater design to the right agreement statistic – Cohen’s and weighted kappa, Fleiss’ kappa, ICC, Krippendorff’s al…
- Net Reclassification Improvement (NRI): The Calculation, and the Case Against It
The NRI compares two risk prediction models by counting who moved in the right direction. This guide gives the arithmetic for both the category-based …
- Exploratory Factor Analysis: Extraction, Rotation, and How Many Factors to Retain
Exploratory factor analysis turns on three decisions usually made by accepting a default: extraction method, number of factors retained, and rotation.…
- Structural Equation Modeling (SEM): Fit Indices, Cut-Offs and Model Evaluation
A judgment-layer guide to evaluating structural equation models: which fit indices to report, why the widely cited Hu and Bentler cut-offs differ from…
- Average Variance Extracted (AVE): Calculation, the 0.50 Threshold, and Discriminant Validity
AVE is the mean proportion of indicator variance a construct explains rather than error. This guide gives the calculation, the origin and conventional…
- Normality of Distribution: How to Check the Normal Distribution Assumption in Research Data
What the normal distribution is, why the Central Limit Theorem makes most explanations of it misleading, and how to actually assess normality using Q-…
- Power Analysis and Sample Size Calculation: A Complete Guide
How to run a power analysis to calculate sample size: the four interlocking quantities (effect size, alpha, power, N), a priori vs. post-hoc power, ch…
- Regression Analysis: Assumptions, Interpretation, and How to Report It
A practical guide to choosing between linear, logistic, and other regression models, checking each model’s assumptions, interpreting coefficients and …
- Cronbach’s Alpha: What It Measures, How to Interpret It, and When to Use Omega Instead
Cronbach’s alpha measures internal consistency, not unidimensionality or reliability in general — the most common misreading of the statistic. This gu…
- Z-Test vs T-Test: Why the t-Test Wins at Every Sample Size, and Where a Z-Test Genuinely Belongs
- Mediator vs. Moderator: Which One Your Hypothesis Actually Needs
Research Tools & Software
- Creating Dummy Variables in Stata: i. Factor Notation vs. tabulate, generate()
Two ways to create dummy/indicator variables in Stata — the i. factor-variable operator vs. tabulate, generate() — and when each one is the right tool…
- Poisson and Negative Binomial Regression in Stata
Fitting poisson and nbreg for count outcomes in Stata: using exposure() to model rates, testing for overdispersion, and reporting results as incidence…
- Logistic Regression in R: glm(), Odds Ratios, and Diagnostics
A runnable R workflow for logistic regression: glm(family = binomial), exponentiating coefficients to odds ratios, why confint() profile intervals dif…
Scholarly Publishing
- Trial Sequential Analysis: Monitoring Boundaries for Cumulative Meta-Analysis
How Trial Sequential Analysis adapts group sequential monitoring boundaries and required information size to cumulative meta-analysis, and what the Co…
Research Information & Identifiers
- OpenAlex: What It Is and How It Works
A practical guide to OpenAlex: the free, CC0 scholarly metadata graph from OurResearch that succeeded Microsoft Academic Graph — its core entities, AP…








