Published Updated

When floating-point calculations change a tiny factorial effect

A 945-case comparison separated effects lost during input rounding from errors introduced by cell means and QR calculations, including cases where centering did not help.

Separating two ways a small effect can change

A tiny difference between experimental groups can disappear when the observations are stored as floating-point numbers. Even when the stored data imply an exact zero effect, a later calculation can produce a nonzero coefficient. We compared those two stages in 945 cases to identify what a future verifier would need to check.

The experiment and its independent reviews were accepted into nomue Protocol's main branch on September 8, 2026 through PR #205. This is bounded engineering research for Release 4 preparation. It selects no production algorithm and adds no supported statistical capability.

A design with an exact reference

We used a complete, balanced two-by-two factorial design: two factors, two levels each, and equal replication in four cells. This small design exposes two main effects and an interaction while allowing exact algebra for the coefficients.

The corpus varies observations per cell from 2 to 16, common offsets of 1, 2²⁰ and 2⁴⁰, seven small effect sizes, and each of the three effect axes. The exact reference uses rational numbers reconstructed from the stored binary64 inputs. Intended rational inputs are retained separately, so input rounding is not mistaken for a later arithmetic error.

We compared CPython builtin-sum cell means, an uncentered QR least-squares calculation, and the same QR route after subtracting the first observation. QR factors a design matrix into orthogonal and triangular parts. The probe calls NumPy's QR operation and then its general linear solver on the triangular factor.

What the 945 cases showed

All exact orthogonality and sums-of-squares partition checks passed. Input projection changed the selected intended coefficient in 666 cases; 525 selected coefficients became exactly zero in the stored inputs.

Calculation routeCases with coefficient errorNonzero result for an exact-zero coefficient
CPython builtin-sum cell means1440
Uncentered QR850462
First-observation-centered QR885501

These are observations for the chosen coefficient in each case, on the disclosed build. Counting a tiny nonzero residual as an error is not a statistical significance test or an overall algorithm ranking.

In one witness, with two observations per cell, offset 2⁴⁰ and intended coefficient 2⁻²⁰, input rounding had already reduced the exact coefficient to zero. The cell-mean route returned zero; uncentered QR returned 0x1.6a09e667f3bcdp-14. Centered QR returned a much smaller but still nonzero value.

In a second witness, with three observations per cell, offset 1 and coefficient 2⁻⁴⁰, the stored inputs preserved the coefficient exactly. Centering did not improve that coefficient's error. Centering is itself a floating-point subtraction, so it must be evaluated as part of the computation rather than assumed to be an exact repair.

Reproducing the comparison

The accepted supplement contains the complete executable Python probe, input construction, exact reference and transcript. The original environment was Linux x86_64, CPython 3.12.13, NumPy 2.3.5 and OpenBLAS 0.3.30. Thread count was not pinned.

The original corpus digest is 2371c1ef31a25816e09d08324718dd87329c5b6aa93076fd17ae772373fff55c. It identifies that execution's output, not a portable required result. The independent review reproduced it on CPython 3.12.11; a later review using a different NumPy/OpenBLAS build obtained different QR rows. CPython 3.11 also changed the direct route because builtin floating summation differs from the disclosed 3.12 behavior.

The initial review identified an inaccurate description of builtin sum as a naive accumulation route. The corrected supplement identifies CPython 3.12's compensated behavior. The close review returned GO with no blocker or should-fix findings and two optional observations about build dependence. Supporting NIST and LAPACK pages were subsequently checked from supplied copies, with provenance limits preserved in the accepted addendum.

What this means for verification

A verification rule needs to distinguish the intended data, the exact mathematical result for admitted inputs, and the result of a particular operation sequence. This experiment gives concrete cases for all three layers as nomue explores factorial inference.

The probe evaluates floating-point coefficients. It does not evaluate floating sums of squares, F statistics, p-values, intervals or significance decisions. It establishes no general error bound, and small normwise QR error need not mean small relative error for an individual near-zero coefficient.

On September 9, 2026, public discussion opened on a bounded Release 4 two-factor specification proposal. That proposal draws on a separate accepted normal-model derivation and later specification reviews. This 945-case study remains coefficient-level evidence; it does not establish numerical support for the proposal. A later SS/F and power-scaling study examines numerical propagation in a separate probe. Unbalanced and rank-deficient designs, general error bounds, and execution admission remain separate work. No Contract, identifier, tolerance, or supported domain was issued by this research acceptance.

Evidence and review

The project records disclose author-side OpenAI assistance and model/provider and context separation for the independent reviews. They do not claim independent human investigators or journal peer review.