Published
What power-of-two scaling can and cannot repair
Exact references show when rescaling recovers a factorial F calculation, and when lost inputs or rounding residuals require a different numerical decision.
A finite answer is only the start of the check
Changing measurement units should leave an F statistic unchanged: both its numerator and denominator change by the same squared scale. A computer can still lose that cancellation if either intermediate number becomes zero or infinity. We extended our 945-case coefficient study to follow sums of squares and F statistics, then examined six small datasets before and after power-of-two normalization.
The numerical research and its corrections have been reviewed and accepted into nomue Protocol. This article explains those bounded experiments. It establishes neither a production algorithm nor a supported range for factorial inference.
Following the calculation beyond a coefficient
The design has two factors, two levels each, and the same number of observations in each of four conditions. A factor sum of squares, SS, measures the variation assigned to an effect. The residual sum of squares, SSE, measures variation left within the conditions. For each of the three effects, the studied F statistic is SS divided by the residual mean square, SSE/df, where df is the residual degrees of freedom.
With n observations per condition, those residual degrees of freedom are 4(n−1). The probes construct exact rational targets from the stored binary64 inputs, separately from their floating-point execution. This distinction identifies whether an error enters when the input is represented, when an intermediate quantity is calculated, or when the final result is rounded.
Across 945 datasets, the extension contains 2,835 effect-specific F targets, of which 2,415 are exactly zero. The largest spurious F for a true zero was about 1.9 × 10−6 for uncentered QR and 7.3 × 10−30 for centered QR; the builtin-cell route produced none on the recorded CPython 3.12 build. Conversely, that route lost 105 nonzero F targets, all below 5 × 10−31. Error counts alone would hide these very different magnitudes. These observations neither rank the algorithms generally nor establish a change in statistical significance.
One transformation, several different outcomes
The follow-up uses six fixtures and two transformations per fixture: identity and normalization by a power of two chosen from the largest absolute input. Three execution routes give 36 floating-point evaluations across the 12 rows.
| Fixture | What normalization showed |
|---|---|
| One dataset at scales 1, 2−600, and 2600 | Lossless rescaling recovers each route’s corresponding unit-scale F outputs. The exact targets are 100, 36, and 4. |
| A common offset of 240 | The transformation loses no inputs, but the routes retain their F errors. It does not repair cancellation caused by the offset. |
| Very large and very small inputs together | A smallest positive subnormal input disappears. The exact F changes, but both the original and transformed targets round to positive infinity. |
| An exact-zero-residual dataset | The intended ratio has a zero denominator. Scaling does not remove spurious positive residuals from the QR calculation. |
At the extreme uniform scales, exact SS and SSE themselves lie outside binary64’s range. Their zero or infinity projections can therefore be correctly rounded, while dividing them produces an invalid 0/0 or infinity/infinity operation. The exact F remains representable. The failure is at the ratio, not necessarily at either projection.
Power-of-two scaling preserves F when the calculation follows the same operations and stays within the necessary floating-point range. The uniform recovery and offset invariance illustrate that conditional property; 36 evaluations are not evidence for every possible implementation or input.
Input loss and a wrong finite answer can have different causes
In the mixed-magnitude fixture, normalization changes the exact F by a relative amount of approximately 2−1674. That is evidence of input loss, but not a materially different binary64 F target: both targets project to infinity. The QR routes instead return a finite value of approximately 2104.9, about 21098 smaller than the exact target.
The explanation lies in a spurious positive SSE. Coefficient rounding leaves nonzero residuals at the six smaller observations; the two observations equal to 0.5 have zero residual. The computed SSE is 0x1.1p-106, while the transformed exact SSE is 2−1204 and rounds to zero. The earlier attribution of these residuals to the two 0.5 observations was corrected during review; the linked successor applies that correction consistently.
A verifier therefore needs separate rules for admitted inputs, the exact target, intermediate calculations, and the final representation. “The result is finite” does not establish agreement with the target.
Reproduction and review
The SS/F corpus was recorded on CPython 3.12 with NumPy 2.3.5; the scaling author run used CPython 3.12.14. The independent scaling review reproduced the rows on CPython 3.12.3, NumPy 2.3.5, and OpenBLAS 0.3.30. CPython summation behavior and NumPy/BLAS builds can change rows. The stored environment and transcript identify an observation, not a portable bitwise requirement.
- Accepted SS/F supplement, with the linked executable probe, exact references, counts, and magnitude qualifications.
- Independent programme review, including the extreme-scale fixture.
- Scaling study with the residual-location correction, linking original and successor scripts and transcripts.
- Scaling review and repair confirmation. The close review also corrects the reviewer’s earlier statements and discloses reuse of the same review context.
The project’s reviews disclose tool assistance and the limits of investigator independence; they are not journal peer review. Release 4 public discussion is open. A later unissued candidate now connects bounded numerical evidence to Record checks, a complete report and controlled execution; that does not make these SS/F probes a validation of any p-value or interval calculation, or establish Protocol support.