Statistical Audit Overview Statistical Validity Score ↳ Normality & Homoscedasticity ↳ Effect Size Reporting ↳ CI Reporting ↳ Power & Sample Size ↳ Multiple Comparison ↳ Replication Type ↳ Missing Data Disclosure ↳ Outlier Pre-Registration Robustness & Sensitivity ↳ Sensitivity Analysis Engine ↳ Conclusion Sensitivity Profile ↳ Multiverse Analysis ↳ Outlier Sensitivity Report Estimation Methods ↳ Bootstrap CI ↳ Permutation Test ↳ Bayes Factor Supplement ↳ Assumption & Robustness Guard Test Configuration Guard ↳ Paired vs Unpaired Guard ↳ Multiple Comparison Enforce. ↳ One-Sided Test Lock ↳ Proportion OLS Prevention Experimental Design Guard ↳ Pseudoreplication Detection ↳ Bio vs Technical Replicate ↳ Batch / Plate Confounding ↳ Repeated Measures Suggestion Confounding & Independence ↳ Independence Formal Check ↳ Confounding Disclosure ↳ Covariate Selection Audit Regression & Modeling Guard ↳ Regression Diagnostics Guard ↳ Compositional Data Warning ↳ Sample Size Justification

Outlier Criteria Pre-Registration

Excluding data points after seeing the results is not outlier removal —it is selective reporting. The legitimacy of any exclusion depends entirely on whether the criterion was declared before the data were analyzed. Without pre-registration of the exclusion criterion, there is no principled way to distinguish a justified exclusion from an attempt to improve a borderline result.

STEP 1 —The Pitfall

The most common sequence in practice: analyze data, observe that one value is extreme, remove it "because it looks like an outlier," reanalyze, observe improvement in the result, report only the clean result. The full sequence is never disclosed. The reported result is a selected result, but nothing in the published figure signals this.

This is a form of p-hacking (detected separately by the p-Hacking Detection feature) that operates at the data exclusion step rather than the test selection step. Its impact on false positive rates is substantial and cumulative across the literature.

[Image placeholder: Two strip charts of the same dataset. Left: with all data points, p=0.12, ns. Right: with one "outlier" removed, p=0.038, *. The removed point is shown as a red dot in both panels to make the manipulation visible.]
Removing a single data point that happens to be in the inconvenient direction changes p=0.12 (ns) to p=0.038 (*). Without pre-declared criteria, this is selective reporting.

STEP 2 —Journal Requirement

ARRIVE 2.0 requires "criteria for including or excluding data," stated before the results section. Nature Methods requires the exclusion criterion to appear in Methods. Increasingly, reviewers ask: "Were the outlier criteria pre-specified?" —and the absence of an answer is treated as a risk of bias.

STEP 3 —Licklider Solution

Input

  • Raw data column selected for analysis
  • Exclusion criterion selection: ROUT (Q=1%), Grubbs test, IQR x1.5 / 3.0, Winsorization, absolute threshold, custom criterion (user-text)
  • Criterion declared before any point is excluded

Output

  • Declared criterion recorded in the Outlier Exclusion Log (Transparency Trail, 1.1.2)
  • Points meeting the criterion highlighted before the user confirms exclusion
  • Three parallel displays: raw data result, excluded-data result, robust estimate (Huber mean or bootstrapped median)
  • If the result changes meaningfully after exclusion (effect size changes by >20%), a disclosure flag is raised
  • Validity Score dimension: Fail if any exclusion occurred without a declared criterion; Warning if criterion declared post-analysis

Guard

The exclusion interface is locked until a criterion is selected. If the user attempts to manually mark a data point for exclusion without first selecting a criterion, the system displays the criterion dialog. Manual exclusions without a criterion are recorded in the log as "manually excluded —no criterion" and automatically set the Validity Score dimension to Warning for advisory review.

[Image placeholder: Licklider three-panel display showing a strip chart with: (1) raw data + result, (2) after ROUT exclusion + result, (3) Huber robust estimate. A disclosure flag shows that the effect size changed by 35% after exclusion.]
Three-panel parallel display: raw, excluded, and robust estimates shown simultaneously so readers can evaluate the impact of exclusion decisions.

STEP 4 —Draft Output (Draft / Needs review)

Outliers were identified using the ROUT method (Q = 1%) applied to the
full dataset prior to inferential analysis. One data point in Group B
met the exclusion criterion (value: 47.2, group mean: 12.3 ツア 2.1 SD)
and was excluded from the primary analysis. Results with all data points
included are reported in Supplementary Figure 1. The exclusion criterion
was declared prior to analysis and recorded in the analysis log.