Published Updated

What a multiple-comparison procedure actually guarantees

Original-paper checks separated overall tests, individual comparisons, and simultaneous intervals, giving future verification rules a more precise statement of what they protect.

What the source checks establish

Multiple-comparison procedures can protect different claims even when they use similar thresholds. Our latest original-paper checks make those differences explicit enough to guide future verification rules: which comparisons belong to the family, which dependence assumptions apply, and which errors a procedure controls.

This extends our six-paper review of multiple-testing procedures with Dunn and Šidák, protected and modified LSD, ordered-range procedures, and the GT2 and Genizi–Hochberg interval constructions. These are public research results for Release 3 preparation. Review and acceptance are recorded by scope; they are not journal peer review or new product support.

The latest GT2/GH source review identified two repairs. Both are recorded in PR #221, which preserves and supersedes the original investigation in #216. September 9 update: the bounded SR-F source acceptance is recorded in PR #224. This supersedes the earlier pending status and does not select the investigated procedures.

A guarantee needs its assumptions

The familywise error rate is the probability of making at least one false rejection in a declared family. Strong control requires the promised bound even when some hypotheses are false and others are true.

For valid individual p-values, Bonferroni's union-bound argument controls that error under arbitrary dependence. Dunn's original planned-comparison intervals provide a particular statistical construction; normality in that construction is not a requirement of the abstract union-bound argument.

Šidák's calibration requires a different joint-probability argument. The inspected paper supports symmetric Gaussian rectangles and specified common-scale constructions, including dependent cases. That does not justify applying the same formula to arbitrary dependent p-values or arbitrary separate standard errors. An implementation needs the actual joint condition, not an unexplained label such as “positive dependence.”

An overall test does not identify the differing pairs

Protected least significant difference, or LSD, first performs an overall ANOVA test and then tests pairs if the overall test rejects. Under the complete null, that first stage limits the probability of proceeding to false pairwise rejections. Strong familywise control is a different question: when some means differ, the first stage can reject while true pairwise equalities still need protection.

The Hayter source investigation separates the protected and modified procedures and their model and design conditions. This matters for a verifier: “ANOVA followed by comparisons” does not specify the pairwise guarantee.

Ordered-range procedures introduce another layer. Their rules can depend on the size of a stretch of ordered means, on containing subsets, and on the ordering of critical values. The Newman, Duncan, Ryan, Einot–Gabriel and Welsch sources therefore cannot be compressed into one interchangeable method name.

The review recorded a printed monotonicity statement alongside finite numerical examples in which raw range quantiles decrease. The cited proof's conditions remain unresolved. Those diagnostics identify a condition that needs checking; they do not prove that the full procedures violate their error guarantees, and they do not authorize silently adding a new monotonicity rule.

Narrower intervals and valid coverage are different questions

GT2 uses the Studentized maximum modulus distribution. The investigated Genizi–Hochberg construction instead uses Studentized range for contrast-preserving transformations and Studentized augmented range for general linear functions. The distributions are different, despite the methods' related purpose.

The source-backed GT2 model has a common unknown variance scale, a known covariance shape, and an independent variance estimate. Unequal estimator variances under that model do not establish a guarantee for arbitrary unknown population variances.

The investigated GH improvement is constrained to two distinct sample sizes and specified width restrictions. A smaller average interval width is not the same as smaller width for every pair. The inspected 1979 corrigendum rejects an inference from failure to attain an optimum within a particular family to Kramer's intervals being liberal. The correction itself does not prove Kramer's coverage.

The review also corrected our own arithmetic interpretation. Averaging the paper's printed rounded components gives 0.998633…, which rounds to its printed 0.9986. Averaging reconstructed, unrounded components gives 0.998662…, which rounds to 0.9987. This difference is not evidence of an error in the source's printed average.

Turning source distinctions into verification requirements

The engineering result is a more precise specification boundary. A future verifier needs the declared family, target, procedure variant, assumptions and output interpretation before it can decide what a successful check means. Correct arithmetic alone cannot supply those missing choices.

The cumulative source set is still incomplete and the Release 3 program remains narrowed. SR-F is accepted as closed for its bounded source-completion scope. Algorithms, numerical constants, tolerances and supported domains require separate decisions. Release 3 public discussion is now open within the supplied-source boundary. Method adoption and numerical support remain separate decisions. The supported scientific scope remains Welch: public local Record verification and the separately announced limited hosted capability.

Evidence and review

The records distinguish original-page inspection, investigator derivation, numerical diagnostics and steward decisions. Authoring and review disclose LLM assistance and separate contexts; repository identities bind the files, while reviewer independence statements retain their stated limits.

Further source work

Three September 9 articles apply these distinctions to unequal-variance comparisons, control and best-treatment targets, and closed-testing graphs. Their source maps and acceptance records are separate from this article’s original snapshot.