Published Updated

Comparing with a control and comparing with the best answer different questions

Source review separates fixed-control tests, step-up and step-down calibration, and intervals that compare each treatment with the best of the others.

Decide what “better” is relative to

Comparing several treatments with a preselected control asks a different question from comparing each treatment with the best of the others. The first fixes the reference before observing results. The second has a target involving an unknown maximum, which the procedure must account for.

Our Release 3 review of Dunnett–Tamhane (1991 and 1992) and Hsu (1984) made those targets and outputs explicit. The bounded source synthesis and its acceptance record have passed review. The practical lesson is to fix the question before choosing a critical value or an interval.

A fixed control still allows different ordered procedures

The inspected many-to-one setting uses normal groups with a common unknown variance. A shared control creates dependence among the treatment comparisons. Accounting for that dependence is part of the procedure, not an optional adjustment after calculating separate t tests.

Source constructionRule to preserveImportant boundary
1991 step-downUse the appropriate joint maximum critical value for the remaining labelled comparisons, stopping at the first non-rejection.Unequal group sizes are allowed in the inspected setting; reordering statistics must also reorder their sample-size and correlation information.
1992 step-upCalibrate critical values through joint ordered-statistic events; the first crossing rejects that comparison and all larger statistics.The bounded characterization uses common correlation and valid, nondecreasing calibration. Substituting step-down constants is not justified.

The adjusted p-value constructions differ too: the step-down construction uses a suffix maximum of local tails, while the step-up construction uses a prefix minimum under the source’s ordering. Local tail probabilities and adjusted p-values cannot be interchanged.

The source review preserves the 1992 paper’s proof and existence limits. Some arguments are referred to an unread technical report, and a larger-family superiority statement remains a conjecture. The accepted result is conditional on the required calibration; it does not certify an unconditional existence theorem for arbitrary family size.

The best of the others is a different target

Hsu’s inspected target is each population location minus the maximum location among all the other treatments. It is neither a difference from a fixed control nor a comparison with a maximum that includes the treatment itself.

The inspected theorem assumes independent, equal-size samples from a common continuous location family, translation-equivariant statistics, and its calibrated-event conditions. A location family shifts the same distributional shape between treatments. In the inspected parametric construction, let Δ compare the observed treatment location with the best other observation, and let d be the model-calibrated width. The interval has the form:

[min(Δ − d, 0), max(Δ + d, 0)]

These intervals include zero. Reading them as ordinary all-pairs intervals, or asking them to guarantee a unique winning treatment, loses the source’s intended selection interpretation. The simultaneous coverage statement is a lower bound under its conditions, not equality at every population configuration.

Checking printed numbers without inventing an erratum

The investigation reconstructed selected table values and checked the meaning of their columns. The independent review used a different numerical construction, including adaptive integration and a cell-count calculation, and checked selected values against printed precision.

Rounded inputs can explain small differences between a reconstructed result and a printed number. That is why the review separates a wording or formula-target problem from evidence of a numerical erratum. It also distinguishes the single-step confidence bounds printed in the 1991 paper from intervals compatible with its step-down testing rule.

The finite diagnostic checks support the recorded examples; they are not general coverage proofs or certified numerical error bounds.

Evidence and what follows

The source reviews disclose prior-context and tool-assistance limits. This article reuses their recorded findings; it adds no independent PDF inspection or numerical run. The accepted public-branch scope retains unread historical dependencies, wider variants, and implementation requirements. These methods remain research characterizations rather than supported nomue operations.

For the related distinction between a family of tests and its guarantee, see What a multiple-comparison procedure actually guarantees.

Release 3 public discussion opened on September 9, 2026. The fixed proposal and evidence map limit the opening to supplied sources and preserve remaining scientific and numerical conditions; these historical findings keep their recorded scope.