Journal Levelize Overview Statistical Disclosure ↳ Error Bar Type Enforcement ↳ Effect Size + CI + N Mandatory ↳ One-Sided / Two-Sided Disclosure ↳ Correlation vs Causation ↳ Causal Language Detection Journal Submission Ready ↳ Nature Reporting Summary ↳ Journal Format Adaptation ↳ Figure Consistency ↳ Accessibility Check ↳ Peer Review Response

Effect Size + CI + N Mandatory

A significance star tells a reader that the null hypothesis was rejected at a given threshold —and nothing more. It does not say how large the effect was, how precisely it was estimated, or how many observations it is based on. Licklider requires all three alongside any significance annotation: effect size, 95% CI, and n per group. Stars-only output is not a Licklider output.

STEP 1 —The Pitfall: Statistical Significance Without Practical Meaning

Statistical significance (p < 0.05) tells you that the observed result is unlikely under the null hypothesis. It does not tell you whether the effect is large enough to matter, whether the estimate is precise, or whether it is based on three or three thousand observations. Two studies can both report p = 0.03 and have completely different implications:

  • Study A: n = 200 per group, Cohen's d = 0.12 (95% CI [0.01, 0.23]). A small, precisely estimated effect. Statistically significant; clinically trivial.
  • Study B: n = 8 per group, Cohen's d = 1.4 (95% CI [0.3, 2.5]). A large, imprecisely estimated effect. Statistically significant; practically important but highly uncertain.

Both studies produce 笘・p < 0.05. Without effect size, CI, and n, the reader cannot distinguish them. This is not a theoretical problem —it is how most figures in published papers currently appear.

STEP 2 —Journal Requirement

The movement away from stars-only reporting has accelerated substantially. Nature requires that "measures of effect size with uncertainty (e.g., 95% CI)" accompany all reported statistical tests. eLife requires "effect size and confidence intervals for all primary results." The American Statistical Association's 2019 statement on p-values and significance states that "we need to move beyond a world of binary thinking ('significant' versus 'not significant')" and recommends routine reporting of effect sizes with confidence intervals.

Over 800 journals have signed the Psychological Science reformulation statement requiring "mandatory reporting of effect sizes and confidence intervals." The American Psychological Association's Publication Manual (7th ed.) requires effect sizes as a standard reporting element. JAMA, BMJ, and The Lancet all include effect size and CI as required elements in their statistical guidelines.

STEP 3 —Licklider's Solution

Input

  • Any figure with significance annotations (star notation or bracket notation).
  • The statistical test result, which provides p-value, test statistic, degrees of freedom, effect size, and 95% CI automatically.
  • N per group, from the N Disclosure system.

Output

  • Mandatory annotation package. Every significance annotation automatically includes: (a) significance level, (b) effect size with type label (Cohen's d, eta2, r, etc.), (c) 95% CI for the effect size, (d) n per group. These are displayed in the figure legend and optionally as hover annotations on the figure itself.
  • Estimation plot suggestion. When stars-only output is attempted, Licklider suggests switching to an estimation plot (Gardner-Altman or Cumming design), which displays raw data, group means, and the effect size with CI on a dedicated difference axis. Estimation plots are available as a first-class figure type.
  • Effect size type auto-selection. The appropriate effect size measure is automatically selected based on the test: Cohen's d for t-tests; eta2 (and partial eta2) for ANOVA; r for non-parametric tests; OR/RR/HR for binary/survival outcomes. The researcher can override the selection with an explicit choice.

Guard

Guard condition —Stars without effect size (blocking): Significance stars or brackets cannot be added to a figure without effect size, 95% CI, and n per group being present in the figure legend. If the statistical test has been run and these values are available, they are automatically added. If they cannot be calculated (e.g., test does not support CI), an explanation is provided and the researcher must acknowledge the limitation before export.

Guard condition —Effect size present but n missing: If effect size and CI are present but n per group is not shown in the legend, export in publication-ready mode is blocked. N per group is the minimum required for a reader to assess whether the effect size is meaningful.

STEP 4 —Draft Output (Draft / Needs review)

Example: Annotation package in figure legend

Statistical comparison: two-sample Welch t-test (two-sided).
**p = 0.003; Cohen's d = 0.91 (95% CI [0.35, 1.47]).
n = 12 (control), n = 12 (treatment).
Full statistical report available in Supplementary Table 1.

Example: Estimation plot legend

Figure 2B. Estimation plot for tumor volume comparison (mm^3, log10 scale).
Left axis: individual data points (gray dots) and group mean ツア 95% CI.
Right axis (difference): mean difference between treatment and control groups
with 95% CI derived by bootstrap (5,000 iterations). Mean difference = 0.42 log-units
(95% CI [0.15, 0.69]); Cohen's d = 0.91 (95% CI [0.35, 1.47]).
n = 12 per group. Two-sample Welch t-test (two-sided): p = 0.003.