Published Updated
Separating format checks from verification results
An experimental Holm verifier separates format and declaration checks from integrity, context and arithmetic results, preserving explicit outcomes for checks that did not run.
Show which kind of check produced the result
Our experimental Holm verifier now reports whether a Record follows its format and declaration rules separately from checks of its integrity, requested context and arithmetic. A caller can inspect those different results without treating a well-formed Record as proof that its calculation was checked.
A Record is the saved analysis and evidence submitted for verification. Conformance means meeting the rules for its structure and declared relationships. Verification checks particular claims about that Record. Both need explicit outcomes: a check that did not run must remain distinguishable from one that passed.
The implementation described below is the historical unissued 0.3.0-candidate.4 in PR #330. It checks ordinary unweighted Holm corrections over supplied p-values. Public verifier support and formal adoption remain unchanged.
September 19 update: the candidate.5 successor now preserves eligible Record checks when caller context differs and can report schema failure when the call is otherwise reportable. Its development checkpoint is merged, not publicly supported. The candidate.4 behavior and evidence below remain the historical record.
Separate the results without inventing another verdict
The preceding candidate placed five checks in one array. The successor places declaration and input-admission results under conformance, alongside evidence that schema admission succeeded. Integrity, equality with the caller’s expected context, and arithmetic results appear under verification_results.
The schema-admission row means that the input passed the format checks needed to produce this candidate report. It does not mean that later semantic checks ran or passed. Schema-invalid input currently produces a refusal, so that row is not a general model for reporting both successful and failed schema checks.
The adapter validates the two sections together. Moving a check to the wrong role, deleting a required result, changing its identity or manufacturing success fails validation. The previous top-level array is rejected in candidate.4. Earlier Records retain their version; the verifier does not silently convert them.
Conformance has no combined verdict, and the report does not declare the research scientifically valid. For an agent consuming the response, the useful evidence is the named check, its outcome and its dependencies.
A new layout does not change which checks can run
Evaluation still proceeds through integrity, caller context, declaration, input admission and arithmetic. The output adapter reconstructs that order when validating the report. Separating the public sections therefore preserves the existing rules for skipped checks and for forwarding the original Record bytes.
For example, a mismatch with the caller’s expected context currently prevents the later declaration checks from running. Their correct result is “not run”; the new section name does not turn them into an independent judgment of the Record. Whether Record-local checks should still run in that case is an explicit question in the public design discussion.
The execution controls also remain in place. An outer execution failure or refusal discards both sections and any forwarded Record bytes. This prevents a partially completed call from retaining an apparently usable fragment of success.
Test misleading combinations as well as valid reports
The author’s retained evidence includes 14 separation controls and nine fixed fixtures with expected outcomes prepared before execution. Existing suites cover 82 envelope controls, 93 public-output controls and 38 budget controls. These are finite check counts, not measurements of agent comprehension or scientific accuracy.
Reproduction uses Node 24.19.0, Python 3.12.14 and the repository lockfile. Actual-host execution records include 33 execution controls, 28 valid public receipts and five original-byte forwards. The numerical sources remain unchanged; these results concern report meaning and execution, rather than an improvement in Holm arithmetic.
An attributed external review reported sound output separation and conditional readiness for bounded reuse. It used different Node and Python versions; envelope and packet snapshot checks did not pass on that environment. Its additional probe scripts and raw logs were not supplied. The fixed review intake retains these limits and does not close final review or adoption.
Make the remaining interpretation choices explicit
Three questions were open at this candidate.4 checkpoint: whether Record-local checks should depend on caller-context agreement; when a format failure can be returned as a failed check rather than a refusal; and how to represent schema admission. The successor's reviewed development design and implementation now address those questions. Formal adoption and the remaining implementation evidence are still separate.
This implementation makes the current distinction inspectable. It advances Licklider’s verification infrastructure by giving callers more precise evidence about what was checked. It does not recompute the supplied raw p-values, prove their statistical validity or demonstrate that an agent will interpret every response correctly.