Published Updated

nomue Protocol opens Release 3 public discussion for independent groups and multiple comparisons

Researchers and developers can comment on how a multi-group analysis should declare its design, comparisons, results, and error-control claims.

Help define what a multi-group verification call should check

Public discussion opened on September 9, 2026 for a nomue Protocol Release 3 proposal covering independent groups and multiple comparisons. Researchers and developers can now challenge how an analysis declares its design, identifies its comparisons, and connects each result to a precise claim.

The proposal brings a catalogue of 49 procedures and a bounded evidence record into one public review. Read the fixed RFC together with its scope and evidence map, then comment on the discussion issue. This is a proposal for future verification rules; method adoption and numerical support remain separate decisions.

Three treatments can produce several different questions

Suppose an experiment compares three treatments using independent experimental units and one continuous measurement per unit. Asking whether the group means differ overall is different from comparing every treatment with a fixed control, comparing all pairs, or examining a contrast chosen after seeing the results.

A contrast is a specified comparison of group means. Its meaning depends on which groups are included and how they are weighted. Testing several claims also raises a separate question: what kind of error should be controlled across the whole family of comparisons?

The RFC asks a future verification call to preserve those choices explicitly. A method name and a p-value cannot supply the missing experimental design or establish a guarantee for a comparison family that was never declared.

  • Design and selection: identify the experimental unit, population, group membership, outcome, assumptions, and when the comparisons were selected.
  • Results and claims: distinguish an overall test, individual comparisons, simultaneous intervals, and different ways of controlling false positives.
  • Procedure and execution: identify the exact variant and define when unsupported inputs, missing declarations, or numerical limitations require refusal.

A visible disposition for every catalogued procedure

The proposal retains all 49 catalogue entries: 15 candidates for Release 3, 27 used as research evidence, five transferred to other work, and two rejected with reasons. Two guidance entries and five explicit catalogue exclusions are recorded separately. These classifications make selection open to scrutiny; they do not announce 49 supported methods.

The current scope is one-way designs with at least three independent groups and one finite continuous outcome per unit. Paired or clustered observations, factorial designs, automatic data preprocessing or method switching, missing-data defaults, and causal claims remain outside this proposal.

The separate Release 4 discussion addresses a bounded two-factor design. The programs have distinct evidence and decision requirements. Release 2 dependencies in the Release 3 proposal remain conditional on their authoritative disposition.

What the opening review established

The source scope is the supplied inventory of 42 numbered originals plus one corrigendum, with later additions allowed through recorded review. That inventory is not a statement that every paper has been fully confirmed or supports every catalogue entry.

The first whole-package review found an incomplete connection between some positive scientific claims and their qualifying independent reviews. The repair explicitly removed those claims from the evidence needed to justify opening discussion, while preserving the candidate classifications and earlier approval records.

The fixed repair review then returned GO for opening readiness, with no blocking or required corrections and one optional link improvement. It confirmed the narrower proposal and reused the earlier package review where applicable. The reviews disclose their assistance and independence limits; this is neither journal peer review nor certification of every method.

In particular, positive claims about Welch calibration, Scheffé coverage, and certain false-discovery-rate procedures remain excluded as scientific support for opening. They can remain questions in the catalogue. Using their guarantees in a later method requires the missing, specifically scoped evidence and review connection.

Our multiple-comparison source work, control-versus-best comparison, and testing-graph analysis explain why these distinctions matter. Their historical findings retain their recorded scope; the RFC's evidence map governs which claims can support the present proposal.

The work between discussion and a supported capability

Scientific and numerical decisions remain explicit. The source ledger still records nine bounded closures, one partial item, and four incomplete items; the overall source set remains incomplete. Opening does not close those items or establish general error-control guarantees.

All candidate numerical support remains held. Earlier claims of rigorous numerical bounds and brackets were withdrawn in PR #265; the affected algorithms are not validated numerical references. Separate work must establish calculation methods, error bounds, tolerances, supported inputs, resource limits, and execution conditions before implementation is promoted.

Current supported scientific scope remains Release 1 Welch, with public local Record verification and the separately announced limited hosted access. This discussion issues no new capability or operational identifier.

September 10: reviewed preparation for candidate development

A limited review of the 15-candidate evidence map found it suitable for progress management, and a separate structural review found the unissued common declarations and result bindings suitable for further exploration. Ordinary Holm is the first planned candidate; its independent-source-review connection and candidate-specific numerical design remain unresolved. These unmerged drafts adopt no method and establish no numerical support. The round handoff records the limits and next steps; the public-discussion scope and window are unchanged.

How to take part

Comment on Issue #274 with the passage concerned, the question or counterexample, and a proposed correction where possible. Feedback on comparison families, selection timing, result meaning, refusal behavior, compatibility, and the evidence boundaries is welcome.

The minimum discussion period is 30 calendar days under the anticipated STABLE-INTENT tier. It began at . The earliest decision is . That date permits consideration of a decision; it is not automatic adoption or a product release date.

Later papers and updates can be added after their identity, inspection, and changed claims are recorded and reviewed. Material changes also require assessment of the discussion scope and window. The opening receipt preserves the fixed inputs and actual start time; the pre-opening wording in the frozen RFC remains part of its history.