Published Updated

What “approximate” means for Games–Howell comparisons

Original-paper checks separate the construction of unequal-variance comparisons from simulation evidence and a guaranteed bound on false positives.

A useful approximation still needs an explicit boundary

Games–Howell compares groups without replacing their different variances with one pooled estimate. That makes it relevant when group variability differs. To turn it into a verification rule, however, we need to know exactly which formula is used and what evidence supports its error rate.

Our Release 3 source review separated Games–Howell, Tamhane T2 and T2′, and Dunnett T3 and C. The bounded source result has passed project review and steward acceptance. These are characterizations for future design, not newly supported nomue methods.

What the comparison protects

The inspected scope is every pair of a fixed set of independent normal groups, with possibly unequal variances and sample sizes. The output under examination is simultaneous two-sided intervals for the true mean differences. “Simultaneous” asks whether all intervals cover their targets together.

For k groups there are k(k−1)/2 pairs. Each pair has its own standard error. Games–Howell and T2 use a Welch-style estimate of degrees of freedom; the table below distinguishes the other critical-value rules. A statement about one interval therefore cannot automatically establish the behavior of the whole family.

ProcedureConstruction that must remain identifiable
Games–HowellA Studentized-range critical value using the pair’s Welch degrees of freedom, multiplied by its standard error and divided by √2.
Tamhane T2A t critical value with a family-size tail adjustment and the pair’s Welch degrees of freedom.
Tamhane T2′A specified modification replacing those degrees of freedom with nᵢ + nⱼ − 2 in stated balance cases. Its simulation rows cannot simply be labelled T2.
Dunnett T3A Studentized maximum-modulus critical value. The reference distribution uses one shared random scale, not independently scaled t variables.
Dunnett CA weighted average of two Studentized-range critical values. Averaging the degrees of freedom instead would define a different calculation.

Even source labels need care: Games–Howell’s “BF” label does not name the Brown–Forsythe procedure called BF in Tamhane’s paper.

A simulated error rate is evidence, not a universal guarantee

The inspected papers document configurations where Games–Howell’s estimated familywise error rate exceeds the nominal 5%. Games and Howell’s Table III includes 0.092 for four groups with sample sizes (11, 8, 4, 3) and variances (1, 3, 5, 7). That is an estimate from the paper’s simulation, not an exact probability for all repetitions or a prediction for a new experiment.

Tamhane’s Table 3 reports joint coverage rather than error rate. Its Games–Howell coverage of 0.916 in an eight-group configuration corresponds to an estimated family error rate of 0.084. Distinguishing coverage from error rate prevents a reversed interpretation.

For the finite-degrees-of-freedom procedures examined here, the papers supply approximate constructions and simulation assessments. Printed T2 or T3 estimates below 5% do not by themselves prove control over every admissible input. Likewise, a reported exceedance does not justify dismissing a method in every setting.

The review does not establish a general finite-sample level-α guarantee for Games–Howell from these assigned texts. It also keeps a cited, unread known-variance proof dependency separate from the finite-sample investigation.

What an implementer should record

A reproducible procedure description needs the comparison family, sidedness, model assumptions, variance estimates, degrees-of-freedom rule, critical-value distribution, and output type. A familiar method name leaves too much of that undecided.

This follows the broader lesson in our earlier multiple-comparison article, now applied to the unequal-variance family. It gives a future verifier something concrete to match: an exact procedure variant with a stated evidence boundary.

Sources and accepted scope

The reviewed originals are Games and Howell (1976), Tamhane (1979), and Dunnett (1980, JASA 75:796–800). The source map gives printed pages, formulas, simulation settings, supplied-file identities, and distinctions between source statements and investigator reasoning.

This article summarizes the recorded source work; it is not an additional independent reading of the PDFs or a new simulation. The reviews disclose their assistance and independence limits. Acceptance covers the bounded source characterization on the public research branch. Release 3 public discussion is now open on the bounded supplied-source proposal. Procedure adoption, numerical support, and one-sided or adjusted-p extensions remain separate decisions.