Published Updated
Cataloguing every multi-group procedure before proposing any of them
We catalogued 49 multi-group comparison procedures, gave each an explicit disposition, and found only seven backed by primary text we had actually read.
Current status
The multi-group semantic research result was merged into the nomue Protocol repository on September 3, 2026, and its source-acquisition follow-up on September 4, each after independent review returned GO. Both are public research input for a possible Release 3 multi-group capability.
September 7 update: a subsequent six-paper primary-source review clarified multiple-testing guarantees and confirmed a conflict between a printed numerical table and its defining equation. The follow-up articleexplains the finding and its bounded source acceptance. The catalogue counts below describe the September 4 baseline. September 9 update: further accepted source work now covers interval guarantees, unequal variances, control and best-treatment comparisons, and testing graphs. These later source results do not retroactively change the historical catalogue counts.
The source acceptances do not select a supported procedure. Release 3 public discussion opened on September 9 on a proposal limited to supplied originals, with unconfirmed positive claims excluded from its opening evidence. No procedure has been adopted or operational identifier, schema, or Public Check issued; the overall source set remains incomplete. Release 1 Welch verification remains the only supported scope.
Why a list of names is not a scope decision
An earlier source review found that a procedure name does not identify what a statistical result guarantees. The meaning depends on the comparison family, the error criterion, the guarantee strength, the assumptions, and the exact variant. See the research note for that finding and its sources.
That leaves an awkward consequence for planning a release. You cannot decide what a multi-group capability will support by listing familiar names, because the names do not carry the properties a check has to verify. The alternative is to write down what replaces the name, apply it to every technique in scope, and see how many survive contact with the actual sources.
Two failure modes make that harder than it sounds. A statistics package's function list quietly becomes the inventory, so whatever the package omits disappears from consideration. And a technique that cannot be sourced gets included anyway, because it is famous and everybody knows what it does.
How the catalogue was built
Inclusion by documented search
The inventory comes from a recorded search rather than from any software catalogue. Thirteen verbatim queries are written into the result, along with backward and forward citation chaining from texts that had already been inspected in full. Every in-scope technique the method finds receives one of four explicit dispositions: implementation candidate, research-only evidence, transfer to a named later program, or reject with a rationale. Techniques ruled out of the design family are listed as recorded exclusions rather than dropped in silence.
An evidence grade on every entry
Each entry carries how well it is known, not just what it is. The grades separate primary text inspected in full, an investigator derivation from such a text, knowledge that exists only through a later text reporting it, and bibliographic identity with no access at all. An entry whose evidence falls short of direct inspection carries the name of the hold that blocks it, and nothing may be frozen on it while that hold is open.
Reuse only inside a recorded scope
Eight primary texts were reused from earlier full-text inspections, with their printed-page pinpoints and artifact digests carried across: Holm (1979), Benjamini and Hochberg (1995), Dunnett (1955), Tukey (1949), Kramer (1956), Hayter (1984), Spjøtvoll and Stoline (1973), and Dunnett (1980). No PDF was re-inspected, and no reused claim was extended to a later variant or a different procedure than the one it was recorded for.
Declarations and refusals designed alongside
The catalogue is paired with what a Record in this family would have to declare rather than let a checker infer: an independent-groups assertion and the experimental unit, explicit per-observation group assignment, a one-way assertion including that the data are not a flattened factorial or clustered design, the admitted observation set, a comparison-family object with its enumerated member set, a selection-timing declaration, a variance model chosen in advance rather than from the observed variances, and a procedure variant identified by an identifier rather than by an eponym. Thirteen fail-closed refusal classes cover what happens when a declaration is missing or the combination is unsupported.
Nineteen counterexample attacks were then run against that design, each recorded with the boundary that rejects it or the deferral it forces.
What the catalogue found
The catalogue holds 49 procedure entries, plus two regulatory framing documents that are not procedures. Of the 49:
- Seven are unblocked implementation candidates. Six rest on directly inspected primary text: Holm step-down, the Holm product-form variant under a declared independence condition, two all-pairs simultaneous-interval procedures, a many-to-one simultaneous-interval procedure, and one false-discovery-rate procedure under an explicit independence declaration. The seventh is Bonferroni, which is unblocked by a derivation from the inspected Holm result rather than by a separately printed single-step guarantee — and is labelled that way instead of being presented as sourced.
- Two are rejected on inspected evidence.
- Two lanes are transferred out with a named destination: rank procedures to the later rank-based program, resampling procedures to the seeded randomness program. Neither is quietly dropped.
- The rest are research-only evidence: catalogued and dispositioned so the inventory is complete, but not proposed for this release.
Cutting across those dispositions, 33 of the 49 entries carry a source-acquisition hold, including eight that would otherwise be implementation candidates. An entry with an open hold keeps its identity and its proposed disposition, and nothing may be frozen on it while the hold stands.
One consequence is worth stating on its own. The proposed dependency order puts the generic p-value family lane before the omnibus lane. Holm-family procedures over a declared comparison family need the design layer plus scalar arithmetic and sorting, while the omnibus F test needs its own primary source and its own numerical tail work. The familiar order — run the analysis of variance first, then follow up — is not the order in which the guarantees can actually be closed.
The attacks fix several boundaries that are easy to lose. An omnibus rejection cannot become a per-pair claim. A many-to-one critical value cannot be reused for an all-pairs family. A balanced-design procedure cannot be extended to unequal group sizes without its own proof. A variance model cannot be selected from the observed variances. Rows cannot be silently removed: an observation is admitted or the check refuses.
Two attacks are recorded as residual risk rather than as solved. A checker working from values cannot detect that the comparison family was chosen after looking at the data, and cannot always detect a factorial design flattened into one-way groups. Both remain declaration-plus-provenance boundaries, with the gap named and tied to an open research hold instead of being claimed as checkable.
Why the scope was narrowed
Four outcomes were available to the investigation: declare the program scope ready, narrow it, defer it, or reject it. It narrowed, and wrote down the reasoning for the choice.
Not ready, because the comprehensive question spans lanes whose primary texts could not be inspected. The objection recorded is about reviewability rather than confidence: a complete disposition ledger whose majority rests on uninspected sources would not be honestly reviewable by anyone else.
Not deferred or rejected either, because a source-established core does exist. It is enough to open bounded work on four lanes and to draft the taxonomy and boundary sections of a future public proposal — on one condition, which the result states explicitly: that the proposal names every hold rather than hiding it.
That condition is the reusable part of this work. Partial evidence is enough to start bounded work, provided the gaps are named in the public artifact instead of being smoothed over in it.
What could not be obtained
A follow-up commission existed to close those holds. As of its reviewed result on 4 September 2026, it had run two passes and closed one of fourteen.
The first pass tried six acquisition routes: direct retrieval over HTTPS, a page-fetch instrument, a general-purpose web index, lawfully supplied local copies, the repository itself, and reachable hosts that turn out to carry no primary text. Of the 48 named sources it acquired none, and it logged the outcome host by host.
The cause is worth being exact about, because it is not the one a reader would guess. Nothing was paywalled out of reach and no publisher declined anything. The investigating environment routes outbound requests through an egress proxy that decides per host whether to open a connection at all, and it declined every scholarly, publisher, archive, library, and government host it was asked for. No encrypted session was ever established, so not even a paywall page was received. The same policy blocked open archives, a preprint server, and a general search engine, which is how the log characterizes it as a policy rather than an access cost: two of the most useful missing papers are free to read on their publisher's own site.
The second pass worked from a supplied packet and inspected three documents in full: a United States regulatory guidance on multiple endpoints and two European regulatory documents. That closed the one hold those documents were assigned to.
What happened next is the part worth reporting. Those three documents describe Bonferroni, Holm, Hochberg, fixed-sequence and gatekeeping procedures, and resampling-based multiplicity adjustment. Treating those descriptions as evidence would have closed several further holds without acquiring anything. The report refuses: they are second-hand descriptions, they are recorded as not used in support of any procedure entry, and the holds stay open.
At that point, around 44 of the required sources remained unread. The result lists each one with the bibliographic identity needed to obtain it, records the refusal per host, and names the cheapest next increment. The remedy it names is ordinary: supply the copies through institutional access, purchase, or library loan, or reach the hosts directly. None of the missing material is lost or unavailable; it is unread, and the catalogue says so on every entry that depends on it.
What the reviews established
The semantic result went through three independent reviews against the exact commit. The first returned GO with no blocker, six should-fix findings, and five nice-to-have findings. A repair answered them, and a close-only review confirmed the closures while leaving one item open. A final repair closed that item and the three findings the repair itself had raised; the final close-only review reported no regression and no new finding of any severity.
The source-acquisition result was reviewed separately, in its own working context, and returned GO with no blocker, should-fix, or nice-to-have finding. That review approved the single hold closure and left the overall input state incomplete.
What this does not establish
This work establishes that the design family can be enumerated honestly, that each entry can carry its evidence grade and its blocking hold, and that the declaration and refusal surface survives nineteen recorded attacks. It selects no procedure and no variant, and it does not make any catalogued technique available.
The complete source set is not ready. The September 4 result closed one of fourteen acquisition holds; the September 7 follow-up records further candidate dispositions separately from formal acceptance. The numerical work for the F distribution, the Studentized range, and multivariate probabilities belongs to a separate commission that is not covered here. No Release 3 identifier, Contract, Profile, schema, bundle, or Public Check exists, and none is proposed by this work.
The bounded Release 3 RFC is now open for public discussion. The fixed proposal and evidence map preserve all 49 dispositions and distinguish candidate questions from supported guarantees. This historical catalogue remains one research input to that discussion. Nothing here changes the current public scope, which remains Release 1 Welch Record verification.
Public evidence
- nomue Protocol pull request #167 — semantic research result, the two repairs, and the review history
- semantic research result — search method, catalogue with evidence grades, declarations, refusal classes, and the nineteen attacks
- nomue Protocol pull request #168 — the three preserved exact-head independent review results
- nomue Protocol pull request #173 — source-acquisition result and its separate independent review
- source-acquisition result — per-host acquisition log, hold dispositions, and the required-source list
- Release 3 preparation record — scope candidate, boundary exclusions, and the conditions for opening public discussion