Published Updated
When a verification call must discard its result
An experimental Holm verifier connects Record checks to shared execution budgets, operating-system limits and cleanup evidence before deciding whether a result can be returned.
A result needs a completed execution behind it
A verification worker can produce plausible output just before its process runs out of memory, exceeds a deadline, or leaves a child process running. Returning that output as an ordinary result would hide an execution failure from the caller.
Our Release 3 Holm candidate now connects Record verification to execution controls that discard results after those failures. The implementation described below, 0.3.0-candidate.3, is an unissued research candidate with retained code, tests and review evidence. It adds no formally adopted Holm support to the public verifier.
The earlier Holm experiment checked whether adjusted p-values belonged to the expected comparisons and inputs. The successor asks an additional question: did the whole call complete under the conditions required to return its evidence?
The successor separates the returned checks
September 19 update: the later candidate.5 development checkpoint now preserves eligible Record-local results across caller-context failures and includes helper-failure repairs. The candidate.3 and candidate.4 evidence below remains fixed historical evidence; it is not substituted for the successor's tests. Actual failed-cleanup and supervisor-loss evidence remains unfinished for that successor.
The historical draft 0.3.0-candidate.4 reports Record conformance separately from integrity, caller-context and arithmetic checks. It retains the private evaluation order, shared budgets, execution controls and original-byte forwarding rules described here. An outer execution failure or refusal still discards both result sections and the forwarded bytes.
The output-separation article explains that report and its additional controls. The fixed candidate.4 source preserves the numerical kernel. At that checkpoint, the final report/refusal boundary and dependency design were unresolved. Candidate.5 revises them in development; formal adoption remains open and Holm support is not newly available.
Keep the checked Record attached to the result
The candidate takes a Record and independently supplied expected context. Its staged checks cover stored integrity, context agreement and the declared supplied-p calculation. The caller cannot establish expected context merely by accepting whatever the submitted Record says.
When all scoped checks pass, the controlled path can forward the original saved Record bytes. It does not reopen the input file after verification, when another process might have changed it. The output adapter also checks that the forwarded bytes match the report’s Record identity and content digest.
A completed call can still report a failed check or a refusal. Completion describes execution; it is not a single overall verdict on the research.
Use one budget for the call, with an outer limit for stalls
The file entry makes two verification passes to preserve the order of input refusals. One inner budget starts before Record file reading and is shared by both passes, so entering the second pass does not reset the clock. Checkpoints continue through worker completion, report validation and final output serialization.
The experimental inner limits are 5,000 milliseconds and 512 MiB of sampled JavaScript heap. Exceeding either replaces eligible output with a processing-limit refusal. Partial checks, payload and verified Record bytes are discarded.
A checkpoint cannot interrupt work between checkpoints. Five seconds is therefore not a hard stopping guarantee, and JavaScript heap measurements exclude Python memory and native allocations. A separate supervisor applies a 30-second outer deadline and Linux cgroup limits to the call’s process group. Those controls cover failures that the inner observations cannot detect in time.
An ordinary process exit cannot override an enforcement event
The supervisor inspects operating-system counters as well as the worker’s exit status. An observed memory-limit or process-count enforcement event invalidates output even if the main process exits with code zero. Cancellation, deadline expiry, output overflow and cleanup failure also prevent result forwarding.
Cleanup runs on successful exits too. The supervisor terminates the call subtree, drains its output, reaps adopted descendants and checks that the call group is empty before removing its resources. Sending a kill request alone is insufficient evidence that cleanup finished.
The private execution receipt retains the observed causes. A separate adapter converts it into a scoped report or refusal. That adapter assumes a locally trusted supervisor receipt; it does not authenticate a receipt supplied by a remote party.
Exercise the failure paths on an actual host
The retained follow-up CI record reports 33 cgroup controls and 28 receipt-projection checks, including five original-byte forwards. An additional test runs an actual Node 22 executable against the Node 24.19.0 candidate: it exits with the reserved unsupported-host code before input access, forwards no result and completes cleanup. These are finite control counts, not workload coverage or a portability guarantee.
Review also repaired a host guard that loaded runtime dependencies too early, restored candidate-identity checks and separated test observations from committed expected snapshots. The active candidate keeps the numerical kernel unchanged. Its result-discard behavior is an execution property, not evidence of improved numerical accuracy.
The public records disclose OpenAI Codex implementation work and the user’s report of joint human and Claude reviews covering PRs #318–#325. Individual reviewer identities and primary-source coverage are not reconstructed where absent. These records do not establish journal peer review or automatically close the remaining research gate.
What remains before public support
The candidate.3 and candidate.4 execution evidence remains preserved at its fixed revisions. Current development continues in candidate.5. Formal adoption still requires matching the final candidate’s claims to applicable reviews, completing implementation evidence, and coordinating identifiers, schemas, checks and conformance expectations under the open RFC.
The candidate checks ordinary unweighted Holm arithmetic over supplied p-values and its declared context. It does not generate those p-values or establish their scientific validity. The engineering result is a narrower, inspectable rule for delivery: a numerical result survives only when the required verification and execution conditions both permit it.