Published

Reducing the cost of historical-analysis reuse decisions with nomue: a paired, equal-evidence LLM evaluation

Adding nomue reduced API cost by 78.1% and task time by 62.9% on historical-analysis reuse decisions the ordinary-tool comparator also answered correctly.

Preprint v1 · not peer reviewed. Published on Zenodo with a supplement for reproducing the reported aggregates.

An earlier analysis can still reproduce its original numbers while no longer being suitable for current use. An agent needs to distinguish the historical computation, current compatibility, method status and permission to use. This study asks whether nomue reduces the work needed to make those decisions from the same evidence.

On the 103 matched pairs where the ordinary-tool comparator also reached the correct decision, adding nomue reduced usage-priced API cost by 78.1% and summed task time by 62.9%. The comparison matches nomue to tasks the comparator solved correctly, excluding comparator failures and incomplete sessions. The full scheduled comparison and a budget-uninterrupted sensitivity analysis also show lower cost and time.

Same task, same evidence

The panel contains 48 synthetic, template-derived Welch-analysis tasks: 16 with resolved current bindings, 16 with rollbacks and 16 with conflicting bindings. Three repetitions per configuration produced 144 scheduled pairs, or 288 sessions. Repetitions do not turn these into 144 independent tasks.

Both configurations used gpt-5.6-sol and received the same historical raw data, archived Welch program, current declarations and policies, plus ordinary Python and file tools. The comparator used those ordinary tools. The nomue configuration additionally required a real call to its historical-analysis recheck capability through MCP.

The outcome was a six-field reuse decision backed by actual execution evidence. This evaluates an assigned workflow; it does not test whether an agent independently chooses to adopt nomue.

Lower cost and time across three comparisons

ComparisonPairsAPI cost reductionTask time reduction
All scheduled pairs14472.9%55.9%
Comparator-correct matched pairs10378.1%62.9%
First two repetitions, both configurations complete9676.7%61.2%

The 103 pairs cover 47 task roots and have correct decisions in both configurations under the frozen task contract. They answer the practical question: when the ordinary-tool workflow succeeds, how much work does nomue save on the same task? This is a supplementary comparison conditional on success, not an estimate over all scheduled attempts.

The 96-pair comparison covers all 48 task roots and was examined after the run to remove the third repetition’s budget interruptions. It is a post-run sensitivity analysis. Among the 103 comparator-correct pairs, nomue had higher API cost on three pairs and was slower on none.

Across the full schedule, usage-priced API cost was $16.490831 for the comparator and $4.475933 with nomue. Costs exclude hosting, labor and tool fees. Time is the sum of individual task durations, not the elapsed duration of the parallel experiment.

Completion, scoring and a comparator-favorable sensitivity

The comparator completed 111 of 144 sessions; nomue completed 144. Conservative cumulative-budget reservations stopped 33 comparator sessions in the third repetition. Those stops are not 33 demonstrated reasoning failures, and the completion difference does not estimate unrestricted decision ability.

Automated six-field scoring counted 99 correct comparator answers and 139 with nomue. Review of answers and execution evidence under the frozen task contract counted 103 and 144, respectively. Correctness here covers those fields and execution evidence, rather than every sentence of the explanation.

Four of the comparator’s eight remaining errors are disputed under a different interpretation of which current source takes precedence. A sensitivity that credits those four comparator answers expands the matched set to 107 pairs and retains reductions of 78.1% in cost and 63.0% in time. It keeps nomue’s frozen-contract scores; it is a comparator-favorable scoring sensitivity, not a symmetric rescoring under an alternative authority rule.

What this establishes

The result supports a concrete workflow benefit: for these historical-analysis reuse decisions, an agent using nomue spent less API budget and task time than the same model assembling the decision with ordinary tools and equal underlying evidence.

The study uses one model and provider, synthetic tasks and intact archives. It does not establish numerical superiority or general reliability across research settings. Diagnostic explanations can remain imperfect even when the scored decision is correct: some conflicts were described as a missing connection rather than a contradiction between active sources.

This product recheck evaluation is separate from the public local nomue Record Verifier MCP package. The paper is evidence about the tested workflow, not an announcement of broader released support. See current product availability.

The author is nomue’s developer and Licklider’s founder. Review included author-led review with AI assistance and AI adversarial review; this is not independent human peer review. The public supplement reproduces aggregates from disclosed rows, not the full private product execution.

Read the paper and reproduce the aggregates

The Zenodo record is licensed CC BY 4.0. Full raw traces, execution workspaces and product internals are not included in the public supplement.

The frozen experiment dates to September 23, 2026 UTC. The manuscript labels its preprint date September 24 in Japan; Zenodo records the publication date as September 23.

Compare the two published evaluations · Earlier study: whole-submission Welch verification