Published Updated

Binding Holm corrections to the intended comparisons

An experiment checks exact Holm adjustments together with the expected declaration and supplied p-values, including changes that leave the displayed answer unchanged.

A correction belongs to a particular set of comparisons

A numerical correction can be internally consistent and still belong to the wrong analysis. For someone checking an agent’s result, the question includes which comparisons were requested, which inputs were supplied, and whether every returned value belongs to that request.

Our Release 3 experiment now connects a caller-selected declaration and supplied p-values to an exact Holm calculation, then checks the submitted context and every adjusted result. The implementation and review records are integrated into nomue Protocol’s public research archive. This is a bounded experiment; it adds no supported public check or product capability.

The successor now checks the whole call

The original connection described here is preserved at its fixed source revision. Subsequent work has assembled the unissued 0.3.0-candidate.3: Record-wide checks, a budget shared across verification passes, operating-system execution controls and rules that discard results after enforcement or cleanup failure.

The execution-control article explains those mechanisms and their actual-host evidence. The numerical kernel remains unchanged; this advances the conditions for returning evidence rather than the accuracy of Holm arithmetic. Formal adoption, public registration and final review applicability remain separate decisions, as recorded in the current release status.

Related reading: separating format checks from verification results.

September 19: source review and a deterministic sorting bound

The successor now retains a bounded numerical review using Holm's original paper. It checks the downstream adjusted-value derivation for significance levels strictly between zero and one, exact decoding and display rounding, ties and evidence comparison. Adjusted p-values are a downstream derivation, not a formula quoted from the paper; supplied raw p-value generation remains outside these checks.

A separate stable merge-sort repair removes dependence on the interpreter's sorting strategy for the comparison-count guard. Its derived worst-case bound is 9,217 comparisons at 1,024 members, below the existing 10,240 guard. The actual all-pairs worker remains limited to 120 members; the larger bound is not an expansion of its admitted scope.

The numerical reviewer also authored that repair. A separate implementation review subsequently checked the bound and worker wiring, with its shared-trust-base and execution limits disclosed. These AI-assisted reviews do not close the whole Research Gate. The latest development checkpoint explains the successor's report behavior and remaining work.

Take the expected context from the caller

The experimental entry point accepts three raw JSON strings: the expected declaration, the expected supplied-p input, and the submitted evidence. The first two establish what the caller intends to check. The submitted evidence cannot choose its own expected analysis.

The connection covers a selected all-pairs family with 3–16 groups and 3–120 comparisons. After input size, structure and declaration-relation checks, it compares the full declaration and supplied inputs using JSON Canonicalization Scheme equality. Object-key order and whitespace can vary; array order remains part of identity.

This binds more than comparison names. Reusing those names with a changed analysis, hypothesis, sidedness, source, revision or other declaration context is rejected. The experiment checks correspondence to the caller’s declaration; it does not establish that the declared experiment actually occurred.

Keep the exact adjusted value and its display separate

The numerical worker treats each supplied binary64 p-value as an exact multiple of 2−1074. It sorts those exact values, applies the Holm rank multipliers, takes successive maxima, caps at one, and maps the adjusted values back to the original comparisons.

The submitted result must match both that exact arithmetic and its nearest, ties-to-even binary64 display. Two different exact adjusted values can round to the same display. A recorded counterexample exercises precisely this case: matching the displayed number alone would miss the changed result.

The connection checks every row. A false final row cannot inherit success from the preceding rows. Context failures stop before the worker starts; an arithmetic mismatch returns no scoped success. A worker failure is recorded as an experiment error rather than treated as an ordinary input refusal.

Test the connection as well as the formula

The retained connection suite records 157 checks in both normal and optimized Python modes, with matching labels and outcomes. These are check counts, not independent datasets. Separate controls cover input substitution, malformed submissions, dependency changes and worker isolation. Small comparison families are also checked against a separately constructed all-subsets Bonferroni reference.

The review led to concrete repairs: honoring the configured Python executable in the integrity suite, isolating the worker from Python path injection, and restoring a dependency alias between tampering controls. A later intake also corrected an overstatement about memory measurements. Live samples and process-reported peaks have different observation windows; the recorded probes do not establish a portable memory bound.

The public records disclose OpenAI Codex authoring, an attributed user-supplied implementation review, and subsequent archive review. Their scopes and provenance are recorded separately. This work is not journal peer review or a new primary-source investigation.

What the result establishes

The accepted outcome is narrowly named declaration_bound_supplied_p_arithmetic_consistent. It establishes agreement with the expected declaration and supplied-p arithmetic under the experiment’s contract. It does not recompute the raw p-values, establish their statistical validity, or return a significance decision. Scientific validity remains explicitly unasserted.

The next adoption decision must address the public contract, registration, supported execution environment and remaining research conditions. The engineering result already makes one useful boundary inspectable: numerical agreement is checked together with the analysis and inputs to which it belongs.