Published Updated
R’s exact Wilcoxon test can return p-values outside the valid range
R’s exact Wilcoxon test returned negative p-values and a value above 1 on a zero-difference input; we reported it with three independent exact-arithmetic checks.
Current status
We filed R Bugzilla PR#19144 on August 28, 2026. The report is open and remains UNCONFIRMED in the Accuracy component.
R Core has not publicly confirmed the cause or accepted a repair. No upstream fix is recorded in the tracker as checked on October 1, 2026.
An upstream commenter did not reproduce the out-of-range symptom with R-devel r90467 on ARM macOS on September 2. That run returned in-range probabilities. In a September 3 follow-up, the reporter compared those posted values with the exact reference and noted a remaining tail accuracy difference. The discussion therefore leaves the out-of-range symptom dependent on the tested build; it does not establish an upstream diagnosis or a repair.
What we observed
A deterministic input of 100 integers, including exactly one zero difference, caused the one-sample exact path of wilcox.test() to return finite p-values outside the valid interval from 0 to 1. The calls completed without a warning and reported the method as Wilcoxon signed rank exact test.
| Alternative | R result | Exact reference in binary64 |
|---|---|---|
two.sided | -7.5495165674510645e-15 | 2.5853625172819263e-18 |
greater | -3.7747582837255322e-15 | 1.2926812586409632e-18 |
less | 1.0000000000000038 | 1.0 |
A p-value is a probability. A negative p-value or a value greater than 1 is outside its mathematical range. For the upper-tail call, the exact p-value is small but strictly positive; it is not zero.
How we checked it
We reproduced the same displayed values in R 4.6.0, R 4.6.1, and an official R-devel build from SVN trunk revision r90451. The R-devel build ran under Rscript --vanilla on x86_64-pc-linux-gnu. Repeated calls were bitwise identical.
Three separately implemented exact-arithmetic lineages agreed on the conditional, zero-preserving signed-rank distribution. R and the exact computations also agreed on the signed-rank statistic, 4755. The disagreement was in the probability calculation, not the observed statistic.
Reporter-side numerical diagnostic
For this input, the floating-point mass function constructed by .dsignrank() summed to 1.0000000000000038. The cumulative probability used by .psignrank() had the same value. Taking its complement therefore produced the negative upper-tail result, and doubling that complement produced the negative two-sided result.
This is a reporter-side numerical diagnostic, not an R Core root-cause attribution. R Core has not publicly confirmed the cause or selected a repair.
Why clamping is insufficient
Restricting the final output to the interval from 0 to 1 would hide the invalid negative value, but it would turn the upper-tail result into zero. The exact upper-tail probability is approximately 1.2927e-18, so that guard would still return the wrong numerical answer.
Summing the floating upper tail directly kept the result inside the valid range, but it remained 26 representable floating-point steps from the correctly rounded exact value. We therefore did not present a final clamp or direct-tail summation as a complete repair, and we did not assume a preferred upstream numerical strategy.
Scope
The report concerns the one-sample exact Wilcoxon signed-rank path with a zero difference in the tested R 4.6.x and R-devel versions. It does not show that every wilcox.test() calculation is affected, and it does not estimate how often ordinary research workflows reach this numerical regime.
R 4.4.3 returned in-range values for the same calls through older behavior that removed the zero and used an asymptotic path. That comparison uses a different calculation path and is not evidence that the new exact result is correct after a simple fallback.
The upstream report remains unconfirmed. This page does not claim an R Core cause, an accepted fix, or availability of a correction in any released R version.
Relation to verification
This case illustrates why a range check and an accuracy check are different. A guard can make a number look valid while discarding a strictly positive tail. A bounded verification result therefore needs to say both which property was checked and what remains unestablished.
The report does not establish that nomue Protocol currently supports Wilcoxon verification, and it is not comparative evidence that nomue improves research-agent performance.
Public evidence
- R Bugzilla PR#19144— current upstream status
- Public report and reproducer— tested versions, deterministic input, exact references, and reporter-side diagnostic
- R-devel Wilcoxon source— official current source for the relevant R-level implementation