Published Updated
Checking factorial statistics without trusting rounded intermediates
Exact arithmetic and probability bounds offer a path beyond scaling repairs, while a review shows why matching rounded answers does not certify an interval.
From finding numerical failures to checking candidate calculations
Our earlier factorial study showed why rescaling data cannot repair every floating-point error. The next step is to calculate exact targets for the stored observations and keep uncertainty explicit when converting those targets into a computer’s output format.
The initial Release 4 candidates separately computed exact arithmetic targets and enclosed upper-tail probabilities for fixed F inputs. Since this article was first published, a successor experiment has connected the raw observations, exact rational F and submitted probability evidence. The candidates and review records are now integrated into the public research archive. They remain experimental and add no supported nomue capability.
This update preserves the original numerical findings below. The new connection article explains the complete-output and evidence-consumer experiments, including the regression for the interval-checker finding.
A rounded average can inflate the residual calculation by 50%
The retained design has two factors, two levels per factor and the same number of observations in each of four cells. Consider three observations in every cell: 1, 1 and the next binary64 number above 1. Let the gap between those two representable numbers be u = 2−52.
The exact cell mean is 1 + u/3. Across all four cells, the exact residual sum of squares, SSE, is 8u²/3. If the mean is first rounded to 1 and residuals are then calculated, SSE becomes 4u²: 50% too large. The stored observations have not changed; the error enters through the intermediate average.
All three exact effects are zero in this example. A reproduced QR calculation nevertheless returns an F statistic of 2.25 for one effect. This is a witness on the recorded environment, not a claim that every QR implementation behaves identically or that a significance decision necessarily changes.
Keep the exact ratio until the representation decision
Every finite binary64 observation is an exact integer multiple of 2−1074. The arithmetic candidate uses that common unit to accumulate integer sums and sums of squares, then represents means and ratios as exact fractions. It avoids calculating residuals from an already rounded mean.
The candidate also distinguishes an exactly zero SSE from a positive SSE that rounds to zero. In the retained proposal, exact zero SSE stops F and tail calculation. A positive denominator does not become mathematically zero merely because its floating-point projection is zero.
The author’s 968 datasets contain 18,380 quantities, all matching a separately implemented rational reference. A subsequent review reconstructed those quantities using pairwise differences and fitted columns, and checked another 137 datasets. These finite tests support the inspected formulas and implementation; they do not select a production environment or establish unrestricted resource limits.
Enclose a small probability instead of subtracting it away
The original tail candidate starts with a fixed binary64 F, treated as an exact number. For the retained degrees of freedom—numerator one and residual 4(n−1)—a change of variable expresses the upper-tail probability as a finite sum of positive terms. This avoids obtaining a tiny tail by subtracting two nearly equal approximate numbers.
Exact fraction arithmetic and a bounded square root provide lower and upper probability bounds. Separate polynomial-integral and remainder-bounded series calculations supply comparison evidence. When both endpoints round to the same binary64 value, that value is determined for every probability inside the interval.
Both routes resolved the same rounded value in 220 cases. In 56 cases a strictly positive mathematical probability rounded to zero. That is different from claiming the probability is exactly zero. High-precision Decimal output was used as a diagnostic, not as proof of an error bound.
Matching rounded answers does not prove an interval is valid
The tail review found an important limit in the research evidence checker. A supplied interval can overlap the oracle interval and produce the correct rounded answer while still excluding the true probability. The reviewer constructed such a replacement interval at n = 2 and F = 4; the checker accepted it.
This does not show a wrong probability in the delivered candidate: the fixed results were reproduced byte for byte, and the enclosure derivation was checked. It shows that the checker’s overlap test is insufficient to certify a supplied interval. Some auxiliary fields and complete-corpus membership are also outside its checks.
The successor evidence consumer now recomputes a fixed candidate enclosure and requires each submitted interval to contain it, with matching endpoint rounding and expected raw-input identity. It rejects a raw-data version of this false-interval witness. This is conditional consistency with the candidate: a tighter valid interval can fail the conservative containment rule, and the runtime does not use the separate test oracle. The original checker and its finding remain preserved in the historical evidence.
The connection is experimental; public adoption remains open
The successor passes exact rational F values into the tail calculation and checks all three effects before accepting a complete result. Its input, representation and work limits are explicit experiment policies. Formal public input limits, execution budgets, supported environments, Record binding and check registration remain adoption decisions.
The public records disclose OpenAI Codex assistance, attributed implementation-review receipts and subsequent archive review; they do not establish journal peer review or closure of the whole research gate. Release 4 public discussion continues independently of these engineering results.
Fixed evidence and reviews
- Successor evidence consumer and implementation-review receipt; archive integration and remaining promotion conditions.
- Arithmetic candidate, formulas and reproduction record; limited arithmetic review.
- Fixed-F tail candidate and source boundaries; tail review and checker counterexamples.
- Preparation-round handoff, preserving completed work, usage limits and remaining conditions.