Evaluation
Measure where nomue makes a difference
Does nomue help an agent reach the right decision? Does it reduce the work needed? What changes in a concrete case? We evaluate these questions through explicit comparisons, inspectable evidence and reproducible analysis.
This page brings that evidence together as the evaluation program grows. Papers, benchmarks and case studies each address part of the question. Start with the dimension that matters to you, then inspect the conditions and materials behind the result.
Decision quality
Checking the whole submission
Can an agent assess a submitted statistical report, including missing claims and conflicts with the requested conditions? The measured outcome is exact correctness of the complete five-field assessment.
216 / 216 expected reports with the bound verifier, compared with 132 / 216 using general Python and 131 / 216 using a stable Welch wrapper.
Tested scope: 72 constructed Welch-report questions, three repetitions and three configurations: 648 sessions using gpt-4.1-mini-2025-04-14.
How to read it: This compares complete configurations on a fixed panel. The verifier’s result path was deterministic. Accepting an alternative output format raises the Python score to 142 / 216. These are complete-report results on the tested tasks, not a general research-accuracy rate.
Study, methods and limitsPaper and reproduction package
Developer-led evaluation · Preprint v1.0 · Not peer reviewed
Cost and time
Reaching a reuse decision with less work
When an agent checks whether an earlier analysis can be reused, does adding nomue reduce API cost and task time? Both configurations received the same underlying historical evidence and ordinary tools; one also required nomue’s recheck capability.
78.1% lower API cost and 62.9% less task timeon the 103 matched pairs the ordinary-tool comparator also answered correctly.
Tested scope: 48 synthetic Welch-analysis tasks, three repetitions and two configurations: 288 scheduled sessions using gpt-5.6-sol.
How to read it: The matched comparison excludes comparator failures and incomplete sessions. Across all 144 scheduled pairs, reductions were 72.9% and 55.9%. Budget reservations stopped 33 comparator sessions; the 96 pairs in the first two repetitions, where both configurations completed every task, showed reductions of 76.7% and 61.2%. Cost is usage-priced API cost; time is summed task duration.
Study, scoring sensitivity and limitsPaper and aggregate reproduction supplement
Developer-led evaluation · Preprint v1 · Not peer reviewed
Concrete cases
See what changes in the answer
Inspect the submitted information, expected assessment and observed outputs side by side. The current collection explains three cases from the whole-submission study.
- A dataset-version conflict despite matching numbers
- A missing standard-error claim
- An incorrect confidence-interval endpoint
Tested scope: Three selected cases, with all three repetitions for each configuration and case IDs that connect the display to the archived evidence.
How to read it: These examples explain observed differences; they are not an additional independent study or an estimate of how often these problems occur. The configurations ran in separate sessions, rather than nomue intervening in another agent’s conversation.
How to assess the evidence
Our aim is to make each claim checkable: what was tested, what counted as success, what the comparator could use, and what another reader can reproduce. As new evaluations are published, they will extend the relevant dimension or add another.
- Defined questions: Identify the task, expected output, metric and tested scope.
- Visible comparisons: State the model, tools, supplied evidence and material differences between configurations.
- Inspectable outcomes: Report denominators, repetitions, interruptions and relevant scoring sensitivities.
- Reproduction materials: Link data and code, and say whether they reproduce saved results or rerun the original workflow.
The current studies provide evidence on specific Welch workflows. They use different tasks, models and endpoints, so their scores are kept separate. They do not yet establish performance across research domains or all released nomue interfaces.
Both studies were led by nomue’s developer. Their individual research pages disclose author interests, review and AI assistance. A published preprint and an archived dataset make evidence available for inspection; they do not constitute independent peer review.
Inspect and reproduce
Start with the public materials for the result you want to check.
- Whole-submission verification package: rescore the 648 saved session rows, inspect panel inputs and expected outputs, and reproduce reported aggregates. The case collection uses this same evidence.
- Historical-analysis reuse supplement: reproduce reported aggregates from 288 session rows, answers, reference fields, case facts, the scoring contract, analysis plan and aggregation script.
These archives support reproduction of the disclosed results. They do not include the private product implementations needed to independently rerun the original nomue sessions. Each study specifies its additional reproduction limits.