Published Updated
What must stay fixed in a multiple-testing graph
Closed testing and graphical procedures make error control inspectable, but order, weights, stopping rules, and zero-level behavior still need precise definitions.
A graph can make a testing plan easier to inspect
A multiple-testing graph assigns error-rate budgets to hypotheses and specifies how rejection can transfer those budgets to other tests. That is a useful way to express a plan. For a computer to execute it faithfully, however, the plan must specify more than arrows and labels.
Our Release 3 source synthesis connects closed testing, fixed sequence, fallback, gatekeeping, and weighted Bonferroni graphs. The bounded source characterization has passed review and steward acceptance. It identifies which assumptions and boundary decisions a future implementation needs to preserve.
The guarantee starts with the family of questions
Closed testing considers the elementary hypotheses and their intersections. An elementary hypothesis can be rejected only when all the required intersections containing it are rejected. If each local intersection test has a valid level-α bound, the intersection of all true nulls supplies the familywise error bound.
That framework does not require independence among already-valid local tests. Each local test still needs its own assumptions. Replacing a Bonferroni local test with a different rule therefore requires separate justification.
| Plan | Meaning to preserve |
|---|---|
| Fixed sequence | Use an order specified in advance and stop at the first non-rejection. Sorting by observed p-values would be a different procedure. |
| Fallback | Keep the predetermined allocations and the rule for carrying or resetting budget after each result. |
| Gatekeeping | Specify the two families, their weights, and exactly which gate condition allows secondary decisions. |
| Weighted Bonferroni graph | Specify nonnegative initial levels, transfer weights and updates, eligible tests, and the stopping rule. |
The inspected graph construction uses a finite, prespecified graph with no self-edges and outgoing weights summing to at most one. Its source relationship to closed testing has a stated scope; it does not make every gatekeeping or closed-testing procedure equivalent to a graph.
Zero can change the pointwise behavior
Consider allocations (0.05, 0) and p-values (0.1, 0). A fixed-sequence procedure stops after the first non-rejection. A rule that continues and literally checks p ≤ level can treat the second, zero-level test differently. Requiring a positive allocated level gives another behavior.
A p-value of exactly zero under a true null has probability zero when marginal validity holds. That fact does not settle what an implementation should do with a false-null zero or a represented floating-point zero. A probability guarantee and the function executed for every input are separate specifications.
Clipping an adjusted value at 1 also needs an endpoint convention. For p = 1 and weight 1/2, the raw ratio is 2 and the clipped value is 1. At α = 1, comparing those two values with α gives different decisions. Neither convention is selected by this research.
A tiny positive value is not the same as a limit
Some graph representations use a positive ε intended to approach zero. The reviewed two-family example shows why taking a small value in code is not automatically the same as the limiting procedure: with ε = 10−9, a sufficiently small secondary p-value can permit a rejection before all the intended gates clear.
The recorded algebra establishes a limit for local weights under the stated positive-weight conditions. It does not establish identical decisions at every finite ε. A verifier must keep the symbolic construction and the executed finite number distinct.
From source reading to an executable rule
The reusable engineering result is a list of decisions that must survive translation into software: family membership, hypotheses, ordering, weights, local tests, allowed α values, zero-level eligibility, transfer updates, and stopping behavior. These should be recoverable from the versioned rule rather than inferred from the observed results.
Review also separated source statements from investigator derivations. For example, the general fallback argument and the sufficient shortcut argument are recorded as author derivations checked in review, not as newly discovered theorems or quotations from an unread reference.
Evidence and scope
The bounded synthesis uses inspected work by Marcus–Peritz–Gabriel (1976), Wiens (2003), Dmitrienko and colleagues (2003), and Bretz and colleagues (2009). A specifically approved source-basis change covers two entries; it does not establish the unread 1995 formulation or historical attribution.
- Part T: six-entry map and numerical boundary examples, especially T.4–T.6.
- Additional verification and source-synthesis review.
- Scoped acceptance and acceptance-record confirmation.
This article summarizes those public research records, including their disclosed assistance and independence limits. It adds no new source inspection. Release 3 public discussion is now open on the bounded supplied-source proposal. Confidence-interval construction, broader variants, and numerical implementation remain separate work. See also our control-versus-best comparison for why output meaning must be fixed alongside a testing rule.