Publications
Engineering
Reproducible numerical checks, implementation work, and upstream contributions to scientific software including SciPy, Boost.Math, R, and Julia.
Upstream reports
Findings and upstream outcomes
14 reports submitted ·5 matching fixes merged ·1 closed without a planned fix
Each report separates the behavior we reproduced, the upstream response, and the status of a repair. A merged fix records code adoption; the report also states whether that fix has reached a release. A proposed repair or a recommended alternative has its own status.
Maintainers weigh practical impact, implementation cost, and available alternatives when deciding whether to change code. A report can therefore close without a repair while retaining its reproducible finding.
SciPy: Student-t: a representable subnormal tail returns zero — Closed as not planned on September 30, 2026 — MPArray recommended; no NumPy-backend fix merged. The report records the rationale and the available alternative. It remains in the submitted-report total and is excluded from the merged-fix total.
Engineering series
How the paired-t evidence was built
Five notes follow the work from a proof pipeline to a reviewed formal-decision packet. Read them in order to see what each stage established and how later review narrowed the remaining questions.
- 1 · Proof pipelineCertifying paired-t numerical evidence before protocol support
- 2 · Floating-point boundariesTesting a paired-t evaluator at floating-point boundaries
- 3 · Tables and per-input checksBuilding paired-t numerical tables and input-specific error checks
- 4 · Full p-value pathTracing a paired-t calculation from observations to p-value
- 5 · Formal decision readinessFrom numerical bounds to a controlled paired-t execution candidate
The Release 2 paired-t candidate now has an independently reviewed formal decision-readiness packet. It assembles the D1–D6 decision ledger, numerical and execution evidence, structural candidates, review dispositions, Release 1 safeguards, and the required coupled landing order. The Steward decisions, authoritative issuance, support activation, and release remain open.
Engineering entries
September 27, 2026
Upstream report
SciPy returns zero for a representable Student-t tail
A Welch test and two lower-level Student-t functions return zero for a positive subnormal probability that remains representable in binary64.
Closed as not planned on September 30, 2026 — MPArray recommended; no NumPy-backend fix merged
September 23, 2026
Upstream report
jStat returns zero for a probability near 44%
Independent integration and a mathematical lower bound expose a noncentral-t probability collapse that remains after an earlier convergence safeguard.
Reported in jStat issue #300 — upstream confirmation pending
September 22, 2026
Upstream report
statsmodels returns zero for a representable chi-square tail
A reported p-value collapse for a representable chi-square tail led to a merged statsmodels repair with regression coverage for related contingency-table paths.
Fix merged into statsmodels main in PR #10276 — not yet released
September 19, 2026
Implementation note
Preserving Welch confidence intervals near zero
A public-source verifier repair retains precision before nearly equal quantities are subtracted, with independent reference checks around the boundary where the repair takes over.
Merged source repair; not in the published npm 0.2.1-rc.1; no Protocol scope or tolerance change
September 19, 2026
Implementation note
Keeping independent checks visible when verification fails
An experimental Holm verifier preserves eligible Record checks when the caller’s expected context differs, and explains why dependent checks cannot run.
Unissued candidate.5 development checkpoint merged; implementation and Research Gate remain open; no additional public support
September 18, 2026
Implementation note
Preserving Student-t probability near zero at one degree of freedom
A closed-form calculation keeps small but representable probability differences from disappearing and gives the public nomue verifier an independent regression check for a known numerical failure.
Implemented in @licklider/nomue-verifier 0.2.1-rc.1; SciPy issue #25667 remains open; no Protocol scope change
September 15, 2026
Implementation note
Making verification results more useful to research agents
nomue development improves difficult calculations, makes completed checks explicit, and preserves the meaning of historical results as versions change.
Product development update — integrated candidates and internal milestones; no new hosted release or public verifier support
September 14, 2026
Upstream report
statsmodels loses finite Welch results when an intermediate sum overflows
Four exactly represented observations per group make statsmodels overflow an intermediate sum, losing finite variances and Welch t-test results.
Fix merged into statsmodels main in PR #10255 — not yet released
September 13, 2026
Implementation note
Separating format checks from verification results
An experimental Holm verifier separates format and declaration checks from integrity, context and arithmetic results, preserving explicit outcomes for checks that did not run.
Historical unissued candidate.4; candidate.5 now implements revised dependencies; formal adoption and support remain open
September 12, 2026
Upstream report
Checking Welch results with exact rescaling
Exact inputs and independent references expose a changed Welch p-value despite a finite result and no warning in a SciPy boundary test.
NumPy float64 repair PR #26209 open; current head is mergeable with 56 / 56 checks passing; issue #26169 open; not merged or released
September 11, 2026
Implementation note
When a verification call must discard its result
An experimental Holm verifier connects Record checks to shared execution budgets, operating-system limits and cleanup evidence before deciding whether a result can be returned.
Historical execution candidate preserved; candidate.5 successor has a merged development checkpoint; no additional supported capability
September 11, 2026
Implementation note
Binding Holm corrections to the intended comparisons
An experiment checks exact Holm adjustments together with the expected declaration and supplied p-values, including changes that leave the displayed answer unchanged.
Original binding experiment preserved; successor includes scoped numerical review and deterministic sort repair; no additional supported capability
September 11, 2026
Implementation note
Checking factorial probability evidence against the raw observations
The factorial candidate now joins Record checks, exact probability evidence, a complete report and controlled execution in an independently reviewed readiness package.
Independently reviewed, unissued Release 4 candidate; D01/D07 amendment discussion open; no Protocol support
September 10, 2026
Technical method
Checking factorial statistics without trusting rounded intermediates
Exact arithmetic and probability bounds offer a path beyond scaling repairs, while a review shows why matching rounded answers does not certify an interval.
Reviewed research components now connected experimentally and archived; no additional supported capability
September 10, 2026
Upstream report
A nonzero Studentized-range tail disappears in SciPy
An exact special-case reference shows SciPy returning zero for a probability near 0.00000002, with no warning in the recorded runs.
Additional reproducer reported — upstream confirmation pending
September 10, 2026
Upstream report
Exact rescaling can reverse SciPy’s Welch ANOVA decision
At an extreme input scale, SciPy’s Welch ANOVA changes a p-value from 0.02650 to 0.05611, crossing the 5% threshold without losing input information.
NumPy float64 repair PR #26209 open; current head is mergeable with 56 / 56 checks passing; issue #26146 open; not merged or released
September 10, 2026
Upstream report
Renaming treatment groups changes agricolae’s REGW result
With observations and group membership unchanged, renaming groups changes a p-value from 0.0363 to 0.0791 and reverses a 5% decision.
Reported by email — upstream confirmation pending
September 9, 2026
Technical method
What power-of-two scaling can and cannot repair
Exact references show when rescaling recovers a factorial F calculation, and when lost inputs or rounding residuals require a different numerical decision.
Reviewed and steward-accepted bounded research; no additional product support
September 9, 2026
Technical method
What “approximate” means for Games–Howell comparisons
Original-paper checks separate the construction of unequal-variance comparisons from simulation evidence and a guaranteed bound on false positives.
Reviewed and steward-accepted bounded research; no additional product support
September 9, 2026
Technical method
Comparing with a control and comparing with the best answer different questions
Source review separates fixed-control tests, step-up and step-down calibration, and intervals that compare each treatment with the best of the others.
Reviewed and steward-accepted bounded research; no additional product support
September 9, 2026
Technical method
What must stay fixed in a multiple-testing graph
Closed testing and graphical procedures make error control inspectable, but order, weights, stopping rules, and zero-level behavior still need precise definitions.
Reviewed and steward-accepted bounded research; no additional product support
September 8, 2026
Technical method
What a multiple-comparison procedure actually guarantees
Original-paper checks separated overall tests, individual comparisons, and simultaneous intervals, giving future verification rules a more precise statement of what they protect.
Source-reviewed Release 3 research; bounded SR-F source acceptance recorded
September 8, 2026
Technical method
When floating-point calculations change a tiny factorial effect
A 945-case comparison separated effects lost during input rounding from errors introduced by cell means and QR calculations, including cases where centering did not help.
Steward-accepted, independently reviewed bounded Release 4 numerical research
September 7, 2026
Technical method
Checking multiple-testing procedures against their original papers
Reviewing six original papers clarified multiple-testing guarantees and exposed a numerical table entry that disagrees with its defining equation.
Independently reviewed SR-C source evidence with bounded acceptance recorded; Release 3 preparation
September 7, 2026
Upstream report
SciPy’s automatic Mann–Whitney U test can change a result when tests are batched
An unchanged sample pair crosses the 5% significance threshold when another pair contains repeated values, because SciPy selects one calculation method for the batch.
Triaged by a SciPy maintainer into scipy.stats; implementation path confirmed — intended behavior and remedy awaiting decision
September 7, 2026
Upstream report
SciPy t-tests can return p=0 or p=1 after exact rescaling
SciPy’s one-sample and paired t-tests can reverse a 5% decision after exact power-of-two rescaling because an intermediate variance underflows or overflows.
PR #24840 merged with the mparray high-precision backend; NumPy float64 repair PR #26135 remains open and unreviewed; issue #26113 open
September 4, 2026
Technical method
Cataloguing every multi-group procedure before proposing any of them
We catalogued 49 multi-group comparison procedures, gave each an explicit disposition, and found only seven backed by primary text we had actually read.
Reviewed catalogue research; bounded Release 3 public discussion now open
September 3, 2026
Upstream report
A Julia signed-rank p-value above 1, now fixed in a release
HypothesisTests.jl returned 1.25 for an exact two-sided signed-rank p-value. The matching correction shipped in v0.12.0 and remains in v0.12.2.
Matching fix released in v0.12.0; present through v0.12.2; issue open
September 1, 2026
Implementation note
From numerical bounds to a controlled paired-t execution candidate
The reviewed Release 2 formal decision packet now assembles the paired-t evidence and required decisions without adopting or issuing Protocol support.
Independently reviewed Release 2 formal decision-readiness package — not adopted, issued, or supported
September 1, 2026
Upstream report
R’s exact Wilcoxon test can return p-values outside the valid range
R’s exact Wilcoxon test returned negative p-values and a value above 1 on a zero-difference input; we reported it with three independent exact-arithmetic checks.
Reported to R — PR#19144 open and unconfirmed
August 31, 2026
Implementation note
Tracing a paired-t calculation from observations to p-value
We connected paired observations to a p-value in one reviewed trace; later work closed its two numerical error ledgers and added a reviewed interval trace.
Independently reviewed p-value and confidence-interval execution traces; interval proof continues
August 31, 2026
Implementation note
Building paired-t numerical tables and input-specific error checks
We built two reviewed 200-value tables and input-specific error checks; later decisions selected the p-value bound and one table for candidate interval work.
Two independently reviewed 200-value tables and input-specific error checks — candidate Release 2 work
August 30, 2026
Implementation note
Testing a paired-t evaluator at floating-point boundaries
We built and independently reviewed a deterministic paired-t probability evaluator and boundary evidence while leaving accuracy bounds, supported inputs, and protocol registration open.
Independently reviewed deterministic evaluator and floating-point boundary evidence — candidate Release 2 work
August 28, 2026
Technical method
Certifying paired-t numerical evidence before protocol support
We built and independently reviewed a proof pipeline for paired-t p-values and critical values before deciding what nomue Protocol will support.
Independently reviewed proof pipeline for paired-t p-values and critical values — candidate Release 2 work
August 26, 2026
Upstream report
A SciPy exact Wilcoxon p-value error, fixed upstream
SciPy’s exact Wilcoxon path could return zero for a positive p-value. SciPy diagnosed the cause and merged a fix the same day it was reported.
Fix merged in SciPy — awaiting a SciPy release
August 23, 2026
Bug report
An extreme-tail sign error in SciPy’s Student-t quantile
A SciPy bug can return positive infinity instead of a large negative value for an extreme-tail Student-t quantile.
Fix merged in Boost.Math — awaiting a Boost release