Publications

Engineering

Reproducible numerical checks, implementation work, and upstream contributions to scientific software including SciPy, Boost.Math, R, and Julia.

Upstream reports

Findings and upstream outcomes

14 reports submitted ·5 matching fixes merged ·1 closed without a planned fix

Each report separates the behavior we reproduced, the upstream response, and the status of a repair. A merged fix records code adoption; the report also states whether that fix has reached a release. A proposed repair or a recommended alternative has its own status.

Maintainers weigh practical impact, implementation cost, and available alternatives when deciding whether to change code. A report can therefore close without a repair while retaining its reproducible finding.

SciPy: Student-t: a representable subnormal tail returns zero — Closed as not planned on September 30, 2026 — MPArray recommended; no NumPy-backend fix merged. The report records the rationale and the available alternative. It remains in the submitted-report total and is excluded from the merged-fix total.

Engineering series

How the paired-t evidence was built

Five notes follow the work from a proof pipeline to a reviewed formal-decision packet. Read them in order to see what each stage established and how later review narrowed the remaining questions.

  1. 1 · Proof pipelineCertifying paired-t numerical evidence before protocol support
  2. 2 · Floating-point boundariesTesting a paired-t evaluator at floating-point boundaries
  3. 3 · Tables and per-input checksBuilding paired-t numerical tables and input-specific error checks
  4. 4 · Full p-value pathTracing a paired-t calculation from observations to p-value
  5. 5 · Formal decision readinessFrom numerical bounds to a controlled paired-t execution candidate

The Release 2 paired-t candidate now has an independently reviewed formal decision-readiness packet. It assembles the D1–D6 decision ledger, numerical and execution evidence, structural candidates, review dispositions, Release 1 safeguards, and the required coupled landing order. The Steward decisions, authoritative issuance, support activation, and release remain open.

Engineering entries

  1. September 27, 2026

    Upstream report

    SciPy returns zero for a representable Student-t tail

    A Welch test and two lower-level Student-t functions return zero for a positive subnormal probability that remains representable in binary64.

    Closed as not planned on September 30, 2026 — MPArray recommended; no NumPy-backend fix merged

  2. September 23, 2026

    Upstream report

    jStat returns zero for a probability near 44%

    Independent integration and a mathematical lower bound expose a noncentral-t probability collapse that remains after an earlier convergence safeguard.

    Reported in jStat issue #300 — upstream confirmation pending

  3. September 22, 2026

    Upstream report

    statsmodels returns zero for a representable chi-square tail

    A reported p-value collapse for a representable chi-square tail led to a merged statsmodels repair with regression coverage for related contingency-table paths.

    Fix merged into statsmodels main in PR #10276 — not yet released

  4. September 19, 2026

    Implementation note

    Preserving Welch confidence intervals near zero

    A public-source verifier repair retains precision before nearly equal quantities are subtracted, with independent reference checks around the boundary where the repair takes over.

    Merged source repair; not in the published npm 0.2.1-rc.1; no Protocol scope or tolerance change

  5. September 19, 2026

    Implementation note

    Keeping independent checks visible when verification fails

    An experimental Holm verifier preserves eligible Record checks when the caller’s expected context differs, and explains why dependent checks cannot run.

    Unissued candidate.5 development checkpoint merged; implementation and Research Gate remain open; no additional public support

  6. September 18, 2026

    Implementation note

    Preserving Student-t probability near zero at one degree of freedom

    A closed-form calculation keeps small but representable probability differences from disappearing and gives the public nomue verifier an independent regression check for a known numerical failure.

    Implemented in @licklider/nomue-verifier 0.2.1-rc.1; SciPy issue #25667 remains open; no Protocol scope change

  7. September 15, 2026

    Implementation note

    Making verification results more useful to research agents

    nomue development improves difficult calculations, makes completed checks explicit, and preserves the meaning of historical results as versions change.

    Product development update — integrated candidates and internal milestones; no new hosted release or public verifier support

  8. September 14, 2026

    Upstream report

    statsmodels loses finite Welch results when an intermediate sum overflows

    Four exactly represented observations per group make statsmodels overflow an intermediate sum, losing finite variances and Welch t-test results.

    Fix merged into statsmodels main in PR #10255 — not yet released

  9. September 13, 2026

    Implementation note

    Separating format checks from verification results

    An experimental Holm verifier separates format and declaration checks from integrity, context and arithmetic results, preserving explicit outcomes for checks that did not run.

    Historical unissued candidate.4; candidate.5 now implements revised dependencies; formal adoption and support remain open

  10. September 12, 2026

    Upstream report

    Checking Welch results with exact rescaling

    Exact inputs and independent references expose a changed Welch p-value despite a finite result and no warning in a SciPy boundary test.

    NumPy float64 repair PR #26209 open; current head is mergeable with 56 / 56 checks passing; issue #26169 open; not merged or released

  11. September 11, 2026

    Implementation note

    When a verification call must discard its result

    An experimental Holm verifier connects Record checks to shared execution budgets, operating-system limits and cleanup evidence before deciding whether a result can be returned.

    Historical execution candidate preserved; candidate.5 successor has a merged development checkpoint; no additional supported capability

  12. September 11, 2026

    Implementation note

    Binding Holm corrections to the intended comparisons

    An experiment checks exact Holm adjustments together with the expected declaration and supplied p-values, including changes that leave the displayed answer unchanged.

    Original binding experiment preserved; successor includes scoped numerical review and deterministic sort repair; no additional supported capability

  13. September 11, 2026

    Implementation note

    Checking factorial probability evidence against the raw observations

    The factorial candidate now joins Record checks, exact probability evidence, a complete report and controlled execution in an independently reviewed readiness package.

    Independently reviewed, unissued Release 4 candidate; D01/D07 amendment discussion open; no Protocol support

  14. September 10, 2026

    Technical method

    Checking factorial statistics without trusting rounded intermediates

    Exact arithmetic and probability bounds offer a path beyond scaling repairs, while a review shows why matching rounded answers does not certify an interval.

    Reviewed research components now connected experimentally and archived; no additional supported capability

  15. September 10, 2026

    Upstream report

    A nonzero Studentized-range tail disappears in SciPy

    An exact special-case reference shows SciPy returning zero for a probability near 0.00000002, with no warning in the recorded runs.

    Additional reproducer reported — upstream confirmation pending

  16. September 10, 2026

    Upstream report

    Exact rescaling can reverse SciPy’s Welch ANOVA decision

    At an extreme input scale, SciPy’s Welch ANOVA changes a p-value from 0.02650 to 0.05611, crossing the 5% threshold without losing input information.

    NumPy float64 repair PR #26209 open; current head is mergeable with 56 / 56 checks passing; issue #26146 open; not merged or released

  17. September 10, 2026

    Upstream report

    Renaming treatment groups changes agricolae’s REGW result

    With observations and group membership unchanged, renaming groups changes a p-value from 0.0363 to 0.0791 and reverses a 5% decision.

    Reported by email — upstream confirmation pending

  18. September 9, 2026

    Technical method

    What power-of-two scaling can and cannot repair

    Exact references show when rescaling recovers a factorial F calculation, and when lost inputs or rounding residuals require a different numerical decision.

    Reviewed and steward-accepted bounded research; no additional product support

  19. September 9, 2026

    Technical method

    What “approximate” means for Games–Howell comparisons

    Original-paper checks separate the construction of unequal-variance comparisons from simulation evidence and a guaranteed bound on false positives.

    Reviewed and steward-accepted bounded research; no additional product support

  20. September 9, 2026

    Technical method

    Comparing with a control and comparing with the best answer different questions

    Source review separates fixed-control tests, step-up and step-down calibration, and intervals that compare each treatment with the best of the others.

    Reviewed and steward-accepted bounded research; no additional product support

  21. September 9, 2026

    Technical method

    What must stay fixed in a multiple-testing graph

    Closed testing and graphical procedures make error control inspectable, but order, weights, stopping rules, and zero-level behavior still need precise definitions.

    Reviewed and steward-accepted bounded research; no additional product support

  22. September 8, 2026

    Technical method

    What a multiple-comparison procedure actually guarantees

    Original-paper checks separated overall tests, individual comparisons, and simultaneous intervals, giving future verification rules a more precise statement of what they protect.

    Source-reviewed Release 3 research; bounded SR-F source acceptance recorded

  23. September 8, 2026

    Technical method

    When floating-point calculations change a tiny factorial effect

    A 945-case comparison separated effects lost during input rounding from errors introduced by cell means and QR calculations, including cases where centering did not help.

    Steward-accepted, independently reviewed bounded Release 4 numerical research

  24. September 7, 2026

    Technical method

    Checking multiple-testing procedures against their original papers

    Reviewing six original papers clarified multiple-testing guarantees and exposed a numerical table entry that disagrees with its defining equation.

    Independently reviewed SR-C source evidence with bounded acceptance recorded; Release 3 preparation

  25. September 7, 2026

    Upstream report

    SciPy’s automatic Mann–Whitney U test can change a result when tests are batched

    An unchanged sample pair crosses the 5% significance threshold when another pair contains repeated values, because SciPy selects one calculation method for the batch.

    Triaged by a SciPy maintainer into scipy.stats; implementation path confirmed — intended behavior and remedy awaiting decision

  26. September 7, 2026

    Upstream report

    SciPy t-tests can return p=0 or p=1 after exact rescaling

    SciPy’s one-sample and paired t-tests can reverse a 5% decision after exact power-of-two rescaling because an intermediate variance underflows or overflows.

    PR #24840 merged with the mparray high-precision backend; NumPy float64 repair PR #26135 remains open and unreviewed; issue #26113 open

  27. September 4, 2026

    Technical method

    Cataloguing every multi-group procedure before proposing any of them

    We catalogued 49 multi-group comparison procedures, gave each an explicit disposition, and found only seven backed by primary text we had actually read.

    Reviewed catalogue research; bounded Release 3 public discussion now open

  28. September 3, 2026

    Upstream report

    A Julia signed-rank p-value above 1, now fixed in a release

    HypothesisTests.jl returned 1.25 for an exact two-sided signed-rank p-value. The matching correction shipped in v0.12.0 and remains in v0.12.2.

    Matching fix released in v0.12.0; present through v0.12.2; issue open

  29. September 1, 2026

    Implementation note

    From numerical bounds to a controlled paired-t execution candidate

    The reviewed Release 2 formal decision packet now assembles the paired-t evidence and required decisions without adopting or issuing Protocol support.

    Independently reviewed Release 2 formal decision-readiness package — not adopted, issued, or supported

  30. September 1, 2026

    Upstream report

    R’s exact Wilcoxon test can return p-values outside the valid range

    R’s exact Wilcoxon test returned negative p-values and a value above 1 on a zero-difference input; we reported it with three independent exact-arithmetic checks.

    Reported to R — PR#19144 open and unconfirmed

  31. August 31, 2026

    Implementation note

    Tracing a paired-t calculation from observations to p-value

    We connected paired observations to a p-value in one reviewed trace; later work closed its two numerical error ledgers and added a reviewed interval trace.

    Independently reviewed p-value and confidence-interval execution traces; interval proof continues

  32. August 31, 2026

    Implementation note

    Building paired-t numerical tables and input-specific error checks

    We built two reviewed 200-value tables and input-specific error checks; later decisions selected the p-value bound and one table for candidate interval work.

    Two independently reviewed 200-value tables and input-specific error checks — candidate Release 2 work

  33. August 30, 2026

    Implementation note

    Testing a paired-t evaluator at floating-point boundaries

    We built and independently reviewed a deterministic paired-t probability evaluator and boundary evidence while leaving accuracy bounds, supported inputs, and protocol registration open.

    Independently reviewed deterministic evaluator and floating-point boundary evidence — candidate Release 2 work

  34. August 28, 2026

    Technical method

    Certifying paired-t numerical evidence before protocol support

    We built and independently reviewed a proof pipeline for paired-t p-values and critical values before deciding what nomue Protocol will support.

    Independently reviewed proof pipeline for paired-t p-values and critical values — candidate Release 2 work

  35. August 26, 2026

    Upstream report

    A SciPy exact Wilcoxon p-value error, fixed upstream

    SciPy’s exact Wilcoxon path could return zero for a positive p-value. SciPy diagnosed the cause and merged a fix the same day it was reported.

    Fix merged in SciPy — awaiting a SciPy release

  36. August 23, 2026

    Bug report

    An extreme-tail sign error in SciPy’s Student-t quantile

    A SciPy bug can return positive infinity instead of a large negative value for an extreme-tail Student-t quantile.

    Fix merged in Boost.Math — awaiting a Boost release