Published Updated
Testing a paired-t evaluator at floating-point boundaries
We built a deterministic Student-t probability evaluator and measured where its binary64 output changes class, without claiming protocol support or a final accuracy bound.
Current status
Two reviewed candidate increments were merged into the nomue Protocol repository on August 30, 2026. The first fixes the operation order of a Student-t probability evaluator. The second measures its output around floating-point class boundaries and records where the implementation and certified mathematical result differ.
This is reviewed candidate engineering evidence, not a supported protocol capability. Release 2 review remains open. No paired-t identifier, Public Check, schema, or bundle has been issued or registered as supported.
Why floating-point boundaries matter
A mathematical probability can be positive while being too small to appear as a positive binary64 number. At the other end, a probability sufficiently close to one is represented as exactly 1. Between those cases are normal and subnormal positive values. A small numerical error near any transition can change the class of the reported result, not just its final digits.
A finite test set can expose these transitions, but its largest observed error is not automatically a guarantee for every input. We therefore separated two tasks: measuring pointwise error at selected boundaries, and proving a future global error bound. Only the first task is complete in this candidate.
The evaluation candidate
The evaluator fixes each binary64 operation and its order. It switches between two positive-series forms at an exact input boundary of |t| = 1. The low-degree paths avoid fragile host-library behavior: the degree-of-freedom 1 path does not call the platform atan function, and the degree-of-freedom 2 tail uses a cancellation-resistant closed form.
The graph also avoids forming t² in the extreme tail, where that intermediate can overflow even when the final probability remains representable. Its iteration cap, branch decision, integer-power order, and refusal conditions are explicit and reproducible.
The reviewed runtime-series evidence contains 19 cases. An independent review route reconstructed the formula and operation graph without using Arb, then compared the TypeScript implementation with an independently written operation mirror over 308 selected inputs. The graph and mirror agreed bit for bit. The measured differences from certified mathematical truth remained observations, not a tolerance or correctness guarantee.
What the boundary evidence found
The follow-up generator searched for adjacent binary64 test statistics on opposite sides of three probability transitions: rounded one to positive normal, positive normal to positive subnormal, and positive subnormal to zero. It used the degree-of-freedom seed 1, 2, 3, 10, 30, 100, and 200. This is a deliberately selected evidence set, not a continuous supported range.
The resulting bundle contains 20 transition cases and 40 endpoints. Across this finite seed, the largest observed graph-to-truth difference was 34 binary64 cells. Ten endpoints placed the graph output and the certified truth in different probability classes. These observations show why a class-boundary rule is needed; they do not establish that 34 is a global maximum.
The candidate records the form of a future safety margin. If a later proof establishes a global error bound B, an otherwise eligible normal or rounded-one result would need to be more than B binary64 cells from the nearest class transition. No value of B has been selected, and this rule is not active at runtime.
Independent review and repair
The runtime-series increment received a candidate-scoped GO review. Its only nice-to-have finding concerned the validator's error message for malformed JSON; that repair was verified in a close-only review.
The first independent review of the boundary-evidence increment returned NO-GO with one blocker and one should-fix finding. The repair strengthened the validator's checks for inverse-beta enclosures, rounding cells, projection classes, precision escalation, dependency identity, and coherent evidence rewrites. A close-only review then confirmed both findings closed and returned GO.
This review history matters because the evidence bundle is intended to fail closed. A matching hash is not enough: the validator must also reject a consistently rewritten bundle whose numerical meaning or maturity claim has changed.
What remains open
The work does not prove a global graph-to-truth error bound, activate a runtime projection margin, or choose a comparison tolerance. It does not select a supported degree-of-freedom maximum, statistic range, numerical domain, platform matrix, runtime inverse-beta table, final table hash, or final refusal-code vocabulary.
The candidate identifiers remain unissued. No authoritative paired-t Public Check, schema, or bundle exists, and Release 2 D5 is not complete. Public review issue #25 remains open until at least September 25, 2026 at 20:52:54 UTC.
This note follows our earlier explanation of how paired-t reference values and critical values are certified. That article covers the proof pipeline; this one covers the executable evaluation graph and its behavior at target-format boundaries.
Subsequent work found a 374-cell pointwise witness outside this finite boundary corpus and built an input-specific error-checking candidate. See the follow-up engineering note for the new candidate tables, evaluator integration, and proof direction.
Public evidence
- nomue Protocol pull request #33 — runtime-series implementation, evidence, repair, and merge record
- runtime-series merge commit
eb4285bf— first merged candidate increment described here - runtime-series review disposition — independent checks, observed numerical behavior, repair, and authority boundary
- nomue Protocol pull request #34 — truth-boundary candidate, review findings, repair, and merge record
- truth-boundary merge commit
6072dd2b— exact merged repository state for the boundary evidence - initial boundary-evidence review — NO-GO disposition and the two findings that required repair
- boundary-evidence close-only review — GO after both findings were closed
- candidate numerical scope — current evidence claims and deliberately open numerical decisions
- Release 2 public review issue #25 — open discussion and future decision record