Covariate Selection Audit
The covariates in your model determine the estimate you report. Selecting covariates after seeing the data —choosing only those that make the primary result significant, or running stepwise selection —is a form of implicit multiple testing. It produces estimates that appear adjusted but are actually optimized for significance. The audit requires that covariate selection be declared and justified before analysis runs.
STEP 1 —The Pitfall
Stepwise regression, backward elimination, and "p < 0.05 to enter" covariate selection strategies are taught in many textbooks as a principled approach to model building. They are not. They systematically inflate Type I error by searching the covariate space for combinations that produce significant treatment effects. A treatment effect that is significant only after stepwise covariate selection is likely to be a false positive.
The subtle version: a researcher runs the primary model without adjustment, observes p=0.06, adds "age" as a covariate, observes p=0.04. Age was chosen not because it is a known confounder but because it produced the desired result. This is covariate fishing, and it is indistinguishable from honest analysis in the reported results.
STEP 2 —Journal Requirement
CONSORT requires that the basis for covariate adjustment be pre-specified in the statistical analysis plan. STROBE requires justification for the covariates included in the primary model. Most high-impact journals now ask: "Were the covariates in the primary model pre-specified?" —and treat stepwise or exploratory covariate selection as a condition for revision or rejection.
STEP 3 — What Licklider Currently Provides
What is described here as methodology guidance
The STEP 1 and STEP 2 sections above describe general methodology for evaluating covariate selection practices. These are standard best practices, not claims about current product automation.
What the product does not do
- The product does not currently record the basis for covariate inclusion (pre-specified vs. data-driven) as a structured field.
- The product does not block stepwise covariate selection.
- The product does not generate model comparison reports (primary model vs. model without data-driven covariates).
- The product does not flag data-driven covariates with advisory text.
Known limitations
This page currently serves as methodology guidance for researchers making covariate decisions. The concepts described in STEP 1 and STEP 2 are important for research quality, but the product does not yet implement the automated checks described in earlier versions of this page.
STEP 4 — Draft Output (Draft / Needs review)
The following is an example of how a researcher might write a covariate selection disclosure, not output that Licklider generates automatically.
The primary analysis adjusted for age, sex, and baseline score —
covariates identified a priori based on their known associations with
the outcome in previous literature (Smith et al., 2021; Jones et al.,
2022) and pre-specified in the statistical analysis plan prior to data
collection. No data-driven covariate selection was performed for the
primary analysis.