Published Updated
SciPy’s automatic Mann–Whitney U test can change a result when tests are batched
An unchanged sample pair crosses the 5% significance threshold when another pair contains repeated values, because SciPy selects one calculation method for the batch.
Same data, different result
Running several statistical tests together can save repetitive work. In the example Licklider reported to SciPy, that change alone moves an unchanged sample pair’s p-value from approximately 0.048 to 0.051. At a 5% significance threshold, the decision changes from rejection to non-rejection.
The trigger is a repeated value in a different sample pair. SciPy’s automatic method selection notices that repetition and switches the calculation method for every pair in the call, including the pair whose data did not change.
Licklider founder Tasuku Kobayashi filed SciPy issue #26115on September 7, 2026. SciPy maintainer and Steering Council member Lucas Colley triaged it into the scipy.stats area. A community contributor then checked current SciPy main and confirmed the global method-selection path and the absence of method='auto' from the existing batched-equivalence test. The issue remains open. Maintainers have not yet decided whether the behavior is intentional or selected a code or documentation change.
The reported example
The Mann–Whitney U test compares two independent samples using their ranks. Here the target samples are [0, 1, 2, 3, 9] and[4, 5, 6, 7, 8, 10, 11]. Their pooled values contain no repetitions, or “ties.” A slice means one pair of samples in a batched call.
| Call | Target p-value |
|---|---|
| Alone, automatic selection | 0.04797979797979798 |
| Batched with another untied pair, automatic selection | 0.04797979797979798 |
| Batched with a pair containing one tie, automatic selection | 0.05131990358807116 |
| Alone, exact method | 0.04797979797979798 |
| Alone, asymptotic method | 0.05131990358807116 |
Every row reports the same target pair, with U = 5. The test is two-sided, continuity correction is enabled, and no multiple-testing correction is applied. Reversing the batch’s row order leaves the target result unchanged. Ties here mean repeated values within one pooled sample pair, not values shared across pairs.
The change comes from method selection
The report records this behavior in SciPy 1.18.1 and development build 2.0.0.dev0+git20260902.fb96fc6, both with NumPy2.3.5. It uses NumPy float64 arrays without NaNs,method='auto', and axis=-1.
In the inspected implementation, when either sample size is at most eight, automatic selection chooses the exact method if there are no ties, and the asymptotic method otherwise. The tie check spans the entire call:_mwu_choose_method(n1, n2, xp.any(t > 1)).
The asymptotic calculation’s tie correction is still computed separately for each slice. Another pair’s ties change the selected method; they do not enter the target pair’s variance correction.
Kobayashi’s mathematical cross-check enumerated all 792 rank allocations for the target pair, giving the inclusive two-sided exact probability38/792 = 19/396. A separate continuity-corrected normal approximation agreed with the asymptotic value. Both numbers match their respective methods in this example. The issue concerns which method is selected.
What the report asks SciPy to clarify
The public discussion now confirms that existing slice-wise equivalence tests cover explicit "exact" and "asymptotic" modes, while the automatic mode is tested separately only with one-dimensional inputs. The contributor asked maintainers to choose between per-slice selection and documenting the current call-wide behavior. Kobayashi replied that per-slice selection would preserve agreement with separate calls, while leaving the intended API behavior for maintainers to decide.
If selecting one method for the whole call is intentional, the report requests documentation explaining that another slice can change a result, along with guidance for users who need agreement with separate calls. If automatic selection should happen per slice, it requests a regression test that mixes tied and untied pairs and compares batched results with separate calls.
Applying the exact method to every pair is not a general workaround: that method does not correct for ties. The report proposes no particular repair and does not measure how frequently the behavior occurs or its effects on false-positive and false-negative rates.
Why this matters for verification
This example shows why a checkable statistical result needs to identify the calculation method actually used, as well as the data and test name. A workflow change from separate calls to a batch can alter automatic choices even when the target observations stay fixed.
For Licklider, the investigation demonstrates reproducible boundary testing, source inspection, and mathematical cross-checks behind our verification work. It does not add Mann–Whitney U support to nomue or establish that nomue detects this behavior. See the current verification scope.
Public evidence
- SciPy issue #26115— Kobayashi’s report, executable reproducer, saved results, complete environment information, maintainer triage into
scipy.stats, and current discussion. - Implementation and test-coverage check— a community contributor confirmed the call-wide selection path on current
mainand identified the missing automatic-mode batch test.