Published Updated
Checking factorial probability evidence against the raw observations
The factorial candidate now joins Record checks, exact probability evidence, a complete report and controlled execution in an independently reviewed readiness package.
Connect the observations to the submitted probability evidence
Our Release 4 experiment now connects stored observations to exact factorial F ratios, bounded upper-tail probabilities and a check of submitted probability intervals. The code and review records are integrated into nomue Protocol’s public research archive. The connection remains experimental and adds no supported public verification capability.
The scope is a balanced, replicated two-by-two design: two factors, two levels each, and an equal number of observations in all four cells. It computes the two main effects and their interaction. For a caller checking a result, the useful change is a single inspectable path from expected raw inputs to all three probability rows.
Pass the exact F ratio to the probability calculation
Our earlier article described separate arithmetic and fixed-F probability candidates. The new connection carries the exact rational F derived from the stored binary64 observations into the tail calculation. Rounding F first would change the input to that calculation.
For numerator degrees of freedom one and residual degrees of freedom 4(n−1), the tail candidate uses a finite sum of positive terms with bounded arithmetic. It returns lower and upper probability endpoints. If both endpoints round to the same binary64 value, they determine the displayed result for any probability inside the enclosure.
The wrapper checks representation and work limits for all three effects before starting tail evaluation. Exactly zero residual sum of squares stops the calculation; a positive residual sum that rounds to zero is refused under this output policy. If a probability remains unresolved within the scheduled work, diagnostics can be retained but the complete result is not accepted.
Require the submitted interval to contain the recomputed enclosure
The consumer receives expected raw cells and a revision independently of the submitted evidence. It checks submission shape, endpoint types and integer-size limits before expensive endpoint arithmetic. Input identity and degrees of freedom must agree before the single recomputation begins.
Each submitted interval must contain the fixed enclosure recomputed by the candidate. Its endpoints must also round to the submitted and expected encodings. All three rows must pass. Overlap with a reference interval, a small numerical difference, or agreement in displayed digits is insufficient.
This is a conservative consistency rule conditional on the candidate enclosure being valid. A mathematically valid interval that is tighter than the recomputed enclosure can be rejected. The runtime does not call the separate probability oracle used in the tests, so acceptance is not an independent proof that the candidate contains the true probability.
A false interval can display the expected answer
The retained regression uses four cells containing [0, 1], [0, 1], [1, 2] and [1, 2]. For the first factor, the exact F is 4 with degrees of freedom (1, 4).
A deliberately false singleton interval lies inside a coarser oracle interval but below a refined lower bound. It therefore excludes the probability while sharing its binary64 display. The new consumer rejects it for failing candidate containment. This turns the earlier overlap-checker finding into an executable regression on raw data.
The reported suite contains 124 consumer checks and 318 wrapper checks, with matching results in normal and optimized Python modes. It includes altered input identity, malformed endpoints, a false final row, tighter valid intervals, unresolved calculations and one-complete-recomputation controls. The wrapper tests use separate arithmetic formulas and a probability oracle.
The successor now enforces worker limits
The controlled execution candidate, integrated into the public research archive through PR #331 wraps the unchanged numerical sources in a controlled worker. Earlier timing measurements described observed runs; this experiment applies operating-system limits and checks how execution ended before exposing a result.
On Linux x86_64 with CPython 3.12.14, one trusted worker receives a 256 MiB virtual-address-space limit, 25-second soft and 26-second hard CPU limits, and a parent-monitored 30-second wall deadline. The memory limit covers that worker’s address space, not total process-tree memory, the supervisor or concurrent calls. Cleanup has a separate two-second wait, and scheduling or kernel stalls prevent a universal return-time guarantee.
The supervisor checks source identity, input and output sizes, worker completion and the complete result structure. Timeouts, allocation failures, crashes, malformed output and cleanup failures suppress result output. Unresolved probability calculations retain a distinct outcome and cannot supply partial successful results.
Review reproduced a cancellation delay when another thread received a termination signal while the main thread waited for output. The repair uses CPython’s signal wakeup descriptor to interrupt that wait. It restores the caller’s signal state and cleans up the worker before propagating cancellation. The new regression passes on the repaired supervisor and fails on its predecessor.
The earlier packet f7be54e, with runtime 1caac8d, remains the fixed record behind the original execution findings. A separate review found no blocking defect in that runtime and identified three low-severity findings. Its reruns used CPython 3.12.3 with the version check patched; that evidence does not validate the supported 3.12.14 environment or independently approve the later repairs.
Check the host and the mode that actually ran
The successor checks that the parent loads the expected modules and that its pipe-signal handling meets the invocation requirements before launching a worker. It refuses incompatible hosts while preserving the caller’s signal setting. The numerical algorithm and worker limits remain unchanged. The review intake and repair record distinguish the original review from these successor changes.
A later check found that running a test driver in optimized Python mode did not propagate that mode to every isolated supervisor process. The earlier statement about normal and optimized coverage therefore needs that qualification. The repaired tests propagate the requested mode and check the mode actually observed inside each isolated process. The production worker retains its separately fixed invocation.
On Linux x86_64 with actual CPython 3.12.14, the recorded CI run on repair commit 761eb8e reports 68 execution controls, ten signal-lifecycle controls, six host-boundary controls and 19 packet controls in each driver mode. It also retains 320 admission rows and seven benchmark calls. These are finite checks of the stated conditions, not a workload success rate or proof of numerical accuracy.
The mode-repair evidence at 2732a26 binds the saved captures to the executed source commit. Its 17-value oracle comparison reads historical captures, not the new run. Those captures are author-run validation. The later candidate-policy record reports the user-supplied T02 CLOSE - GO disposition for this packet; that bounded completion does not close numerical-method review or public adoption.
Make saved evidence fail when its checks fail
The execution candidate and its evidence checks were integrated into the public research archive on September 14, 2026. The follow-up in PR #334 checks the saved evidence against an immutable commit, including the complete file inventory, source bindings and hashes. Changing a saved result and its declared hash together does not bypass that anchor.
The original oracle could report disagreement or empty coverage while returning a successful process exit. A new adapter makes numerical disagreement, missing or empty comparisons, duplicate observations and non-finite values fail the CI step. The original numerical sources, historical captures and comparison tolerance remain unchanged.
Thirteen validation tests passed in each normal and optimized Python mode; the unchanged historical comparison covers 17 values. The preservation and review record explains the boundary: these checks establish archive consistency, not authenticity of the original CI output or independent mathematical clearance. The oracle reads historical captures, not each new execution.
Let the verifier generate the evidence needed for a comparison
The selected policy for the next public candidate requires the producer to declare the p-value, with numerical evidence generated and checked by the verifier. It does not require the producer to submit the probability intervals used in the historical consumer experiment above. That experiment remains evidence about a fixed consistency check, rather than the required public Record format.
The policy also distinguishes a proved mismatch from an undecided comparison. If the evidence rules out a declared value, the check can fail even when another quantity remains undecided. A check passes only when every required comparison passes. A timeout or failed execution does not become an undecided numerical result, and a positive probability that rounds to zero cannot match a declared zero.
These are recorded candidate policy choices, not behavior newly available in the public verifier. Subsequent work selected concrete candidate procedures, bounded inputs and execution conditions, report and command-line behavior, and review evidence. Formal adoption, authoritative registration, public CLI treatment and support activation remain separate work.
Join the numerical path to a complete candidate report
The T09–T14 candidate chain connects Record admission, the exact factorial ratios and bounded tail calculations described above, comparison results, a complete report and controlled execution. The integrated candidate and its checks entered the Protocol research archive through PR #348.
The final-readiness intake records a GO disposition for the fixed candidate. The adoption-readiness packet, preserved through PR #350, records that the remaining A1 review findings were closed without a new finding.
The public intake discloses that it preserves maintainer-provided review results rather than every original reviewer receipt. The result is therefore a reviewed, unissued candidate ready for a formal adoption decision—not an issued Requirement, public verifier feature or supported Release 4 capability.
Keep two proposed comparison rules on their own clock
A separate amendment discussion now covers D01 and D07. D01 first projects both values to binary64 with nearest-ties-to-even rounding, then requires strict equality. D07 distinguishes a completed indeterminate comparison from both a pass and a proved mismatch, without fabricating a point result.
The fixed amendment input does not change the existing five CLI exit-code meanings. Its controlling 30-day window starts from GitHub's comment creation time, , with an earliest decision on . This corrects the previously reported body-declared time of 01:50:49 UTC. It does not reset the original proposal’s October 9 clock, and neither date automatically adopts the candidate.
The formal decision packet and close-only repair confirmation are now merged. Packet review is complete; steward decisions and the coupled Protocol/verifier integration remain open. A completed indeterminate-only report cannot use exit code zero under the existing meanings: the required decision must constrain the supported procedure to resolved outcomes or defer that report state to a separately versioned successor.
The remaining boundary is explicit
The implementation admits only bounded inputs and operand sizes. Its outer range of 2–65 observations per cell is not a promise that every dataset in that range passes: admission also depends on representation and workload. Recorded resource probes describe their hosts and cases, not portable latency or memory guarantees.
The public evidence discloses OpenAI Codex authoring, attributed user-supplied review results and subsequent archive review. These records support bounded implementation claims; they do not establish scientific model validity. The successor candidate now binds a Record to a complete report and selects bounded inputs, execution conditions and limits for review. Formal registration, public CLI treatment, support activation and Release 4 adoption remain open.