Evaluation cases
What a plausible report can miss
A correct calculation is only part of checking a submitted result. These three cases show where the complete assessment changed when the reporting conditions, a required claim, or one interval endpoint was checked.
From a developer-led study of 72 constructed Welch-report questions, using gpt-4.1-mini-2025-04-14. B, W, and C were separate sessions on the same inputs, each repeated three times. They are not a transcript in which nomue intercepted an error made by B or W.
Case 01 · s16_condition_label_01
The numbers matched. The dataset version did not.
Submitted report
The requested analysis named dataset version fresh-panel-8-0-v2. The submitted report named fresh-panel-8-0-v1. Its reported Welch numbers matched the reference for the supplied data.
Expected assessment
The correct five-field report records a dataset_version conflict and marks the overall submission MISMATCH, even though its numerical field is MATCH.
Observed overall verdicts
| Configuration | Run 1 | Run 2 | Run 3 |
|---|---|---|---|
| B · General Python | MATCH | MISMATCH | MATCH |
| W · Stable wrapper | MATCH | MATCH | MISMATCH |
| C · Bound verifier | MISMATCH | MISMATCH | MISMATCH |
In the first run, both B and W explicitly listed the version conflict but still returned overall: MATCH. C listed the same conflict and returned overall: MISMATCH. The other repetitions show that B and W could also get the overall verdict right.
Case 02 · s16_missing_01
A missing standard error changed the answer.
Submitted report
The submitted Welch report omitted mean_difference_standard_error. Other numerical claims were present.
Expected assessment
The expected report names the missing field and returns UNVERIFIABLE for both the numerical and overall assessment. It does not fill in the absent claim from a fresh calculation.
Observed overall verdicts
| Configuration | Run 1 | Run 2 | Run 3 |
|---|---|---|---|
| B · General Python | MATCH | No report | MATCH |
| W · Stable wrapper | No report | No report | No report |
| C · Bound verifier | UNVERIFIABLE | UNVERIFIABLE | UNVERIFIABLE |
In two B runs, the final report listed no missing fields and returned MATCH. C named the absent standard error in each run. “No report” means the run did not yield an accepted five-field report; it is not an incorrect MATCH verdict. The study did not give B and W the complete canonical field vocabulary, which limits how their behavior on missing-field cases should be interpreted.
Case 03 · s16_single_04
One confidence-interval endpoint escaped review.
Submitted report
The submitted Welch report included an incorrect upper endpoint for its 95% confidence interval while the other reported fields matched the reference.
Expected assessment
The expected report identifies mean_difference_ci_high as the single mismatched field and returns MISMATCH for the numerical and overall assessment.
Observed overall verdicts
| Configuration | Run 1 | Run 2 | Run 3 |
|---|---|---|---|
| B · General Python | MATCH | MATCH | MISMATCH |
| W · Stable wrapper | MATCH | MATCH | No report |
| C · Bound verifier | MISMATCH | MISMATCH | MISMATCH |
In the first two runs, B and W accepted the submission as MATCH. C identified the upper endpoint in all three runs. B got the complete report right on its third run; W produced no accepted report on that run.
Inspect the evidence
These are selected illustrations from the published fixed panel, not a sample of everyday research submissions or an estimate of real-world error rates. B used general Python, W added a stable Welch calculation wrapper, and C used a bound verifier that returned the completed five-field assessment. Binding, report construction, and the structured return changed together in C.
The original study scored complete five-field reports: B returned 132/216, W returned 131/216, and C returned 216/216 expected reports. C's result path was deterministic and confirmed before the scored run. The comparison does not isolate a binding-only effect or establish nomue's accuracy on real research submissions, other methods, or other models.
To inspect these cases, download the replication package from Zenodo and search their case IDs in INPUTS.json, EXPECTED.json, and the 648 rows of RESULT.sanitized.json. Runpython reproduce.py from the extracted package to rescore every saved report. ANSWERS.json retains final messages needed for the study's output-format sensitivity analysis.
The public archive supports rescoring saved evidence. It does not contain the private C verifier or complete provider transcripts, so it does not independently rerun the original sessions. These cases are not public nomue Records for the local Release 1 verifier.
Download the source packageRead the study design and limitsExplore all evaluation evidenceExplore current nomue capabilities