Evaluation cases

What a plausible report can miss

A correct calculation is only part of checking a submitted result. These three cases show where the complete assessment changed when the reporting conditions, a required claim, or one interval endpoint was checked.

From a developer-led study of 72 constructed Welch-report questions, using gpt-4.1-mini-2025-04-14. B, W, and C were separate sessions on the same inputs, each repeated three times. They are not a transcript in which nomue intercepted an error made by B or W.

Case 01 · s16_condition_label_01

The numbers matched. The dataset version did not.

Submitted report

The requested analysis named dataset version fresh-panel-8-0-v2. The submitted report named fresh-panel-8-0-v1. Its reported Welch numbers matched the reference for the supplied data.

Expected assessment

The correct five-field report records a dataset_version conflict and marks the overall submission MISMATCH, even though its numerical field is MATCH.

Observed overall verdicts

ConfigurationRun 1Run 2Run 3
B · General PythonMATCHMISMATCHMATCH
W · Stable wrapperMATCHMATCHMISMATCH
C · Bound verifierMISMATCHMISMATCHMISMATCH

In the first run, both B and W explicitly listed the version conflict but still returned overall: MATCH. C listed the same conflict and returned overall: MISMATCH. The other repetitions show that B and W could also get the overall verdict right.

Case 02 · s16_missing_01

A missing standard error changed the answer.

Submitted report

The submitted Welch report omitted mean_difference_standard_error. Other numerical claims were present.

Expected assessment

The expected report names the missing field and returns UNVERIFIABLE for both the numerical and overall assessment. It does not fill in the absent claim from a fresh calculation.

Observed overall verdicts

ConfigurationRun 1Run 2Run 3
B · General PythonMATCHNo reportMATCH
W · Stable wrapperNo reportNo reportNo report
C · Bound verifierUNVERIFIABLEUNVERIFIABLEUNVERIFIABLE

In two B runs, the final report listed no missing fields and returned MATCH. C named the absent standard error in each run. “No report” means the run did not yield an accepted five-field report; it is not an incorrect MATCH verdict. The study did not give B and W the complete canonical field vocabulary, which limits how their behavior on missing-field cases should be interpreted.

Case 03 · s16_single_04

One confidence-interval endpoint escaped review.

Submitted report

The submitted Welch report included an incorrect upper endpoint for its 95% confidence interval while the other reported fields matched the reference.

Expected assessment

The expected report identifies mean_difference_ci_high as the single mismatched field and returns MISMATCH for the numerical and overall assessment.

Observed overall verdicts

ConfigurationRun 1Run 2Run 3
B · General PythonMATCHMATCHMISMATCH
W · Stable wrapperMATCHMATCHNo report
C · Bound verifierMISMATCHMISMATCHMISMATCH

In the first two runs, B and W accepted the submission as MATCH. C identified the upper endpoint in all three runs. B got the complete report right on its third run; W produced no accepted report on that run.

Inspect the evidence

These are selected illustrations from the published fixed panel, not a sample of everyday research submissions or an estimate of real-world error rates. B used general Python, W added a stable Welch calculation wrapper, and C used a bound verifier that returned the completed five-field assessment. Binding, report construction, and the structured return changed together in C.

The original study scored complete five-field reports: B returned 132/216, W returned 131/216, and C returned 216/216 expected reports. C's result path was deterministic and confirmed before the scored run. The comparison does not isolate a binding-only effect or establish nomue's accuracy on real research submissions, other methods, or other models.

To inspect these cases, download the replication package from Zenodo and search their case IDs in INPUTS.json, EXPECTED.json, and the 648 rows of RESULT.sanitized.json. Runpython reproduce.py from the extracted package to rescore every saved report. ANSWERS.json retains final messages needed for the study's output-format sensitivity analysis.

The public archive supports rescoring saved evidence. It does not contain the private C verifier or complete provider transcripts, so it does not independently rerun the original sessions. These cases are not public nomue Records for the local Release 1 verifier.

Download the source packageRead the study design and limitsExplore all evaluation evidenceExplore current nomue capabilities