Statistical Audit Overview Statistical Validity Score ↳ Normality & Homoscedasticity ↳ Effect Size Reporting ↳ CI Reporting ↳ Power & Sample Size ↳ Multiple Comparison ↳ Replication Type ↳ Missing Data Disclosure ↳ Outlier Pre-Registration Robustness & Sensitivity ↳ Sensitivity Analysis Engine ↳ Conclusion Sensitivity Profile ↳ Multiverse Analysis ↳ Outlier Sensitivity Report Estimation Methods ↳ Bootstrap CI ↳ Permutation Test ↳ Bayes Factor Supplement ↳ Assumption & Robustness Guard Test Configuration Guard ↳ Paired vs Unpaired Guard ↳ Multiple Comparison Enforce. ↳ One-Sided Test Lock ↳ Proportion OLS Prevention Experimental Design Guard ↳ Pseudoreplication Detection ↳ Bio vs Technical Replicate ↳ Batch / Plate Confounding ↳ Repeated Measures Suggestion Confounding & Independence ↳ Independence Formal Check ↳ Confounding Disclosure ↳ Covariate Selection Audit Regression & Modeling Guard ↳ Regression Diagnostics Guard ↳ Compositional Data Warning ↳ Sample Size Justification

Covariate Selection Audit

The covariates in your model determine the estimate you report. Selecting covariates after seeing the data —choosing only those that make the primary result significant, or running stepwise selection —is a form of implicit multiple testing. It produces estimates that appear adjusted but are actually optimized for significance. The audit requires that covariate selection be declared and justified before analysis runs.

STEP 1 —The Pitfall

Stepwise regression, backward elimination, and "p < 0.05 to enter" covariate selection strategies are taught in many textbooks as a principled approach to model building. They are not. They systematically inflate Type I error by searching the covariate space for combinations that produce significant treatment effects. A treatment effect that is significant only after stepwise covariate selection is likely to be a false positive.

The subtle version: a researcher runs the primary model without adjustment, observes p=0.06, adds "age" as a covariate, observes p=0.04. Age was chosen not because it is a known confounder but because it produced the desired result. This is covariate fishing, and it is indistinguishable from honest analysis in the reported results.

[Image placeholder: Simulation showing distribution of p-values when: (1) covariate pre-declared xuniform p-value distribution under null (correct). (2) Best covariate selected from 10 candidates after seeing data xp-value distribution skewed toward 0 (inflated Type I error). Shows that 10 covariates tried and the best one selected gives an effective Type I error of ~40% at the nominal alpha=0.05 threshold.]
Type I error inflation from data-driven covariate selection: 10 covariates tested, best selected xeffective alpha ~40% at nominal alpha = 0.05.

STEP 2 —Journal Requirement

CONSORT requires that the basis for covariate adjustment be pre-specified in the statistical analysis plan. STROBE requires justification for the covariates included in the primary model. Most high-impact journals now ask: "Were the covariates in the primary model pre-specified?" —and treat stepwise or exploratory covariate selection as a condition for revision or rejection.

STEP 3 — What Licklider Currently Provides

What is described here as methodology guidance

The STEP 1 and STEP 2 sections above describe general methodology for evaluating covariate selection practices. These are standard best practices, not claims about current product automation.

What the product does not do

  • The product does not currently record the basis for covariate inclusion (pre-specified vs. data-driven) as a structured field.
  • The product does not block stepwise covariate selection.
  • The product does not generate model comparison reports (primary model vs. model without data-driven covariates).
  • The product does not flag data-driven covariates with advisory text.

Known limitations

This page currently serves as methodology guidance for researchers making covariate decisions. The concepts described in STEP 1 and STEP 2 are important for research quality, but the product does not yet implement the automated checks described in earlier versions of this page.

[Image placeholder — not current product UI. Illustrative concept of a covariate audit panel.]
Illustrative concept only. The current product does not expose a dedicated covariate audit panel.

STEP 4 — Draft Output (Draft / Needs review)

The following is an example of how a researcher might write a covariate selection disclosure, not output that Licklider generates automatically.

The primary analysis adjusted for age, sex, and baseline score —
covariates identified a priori based on their known associations with
the outcome in previous literature (Smith et al., 2021; Jones et al.,
2022) and pre-specified in the statistical analysis plan prior to data
collection. No data-driven covariate selection was performed for the
primary analysis.