Roadmap
From one proved verification call to shared research infrastructure
Welch is the first working, publicly checkable vertical slice, not the product boundary. The adopted product sequence expands both the platform beneath each call and the scientific methods available through it.
The stages below separate public artifacts, release candidates, next release work, public discussions, and adopted plans. Protocol proposals advance on their own evidence; Release 3, Release 4 and Release 5 discussions are open with distinct scopes and evidence requirements. These stages do not promise product ship dates.
Product development since limited Release 1
Since limited Release 1, nomue development has expanded candidate calculation range, improved completion of difficult calculations, clarified what was checked, and added historical-result handling.
Development update: integrated candidates and completed internal milestones are not a new hosted release or an expansion of public verifier support.
Available now
A public trust layer for Welch verification
Anyone can install the public @licklider/nomue-verifier package from npm and run nomue verify locally to check a conforming Release 1 Record for independent two-group continuous outcomes under the two-sided Welch two-sample t procedure. It recomputes the covered numerical quantities and returns a machine-readable report of the scoped checks without calling a nomue server after installation.
nomue Protocol Release 1 Public Draft
Public Record, bundle, check, and result meaning for the first supported scope.
@licklider/nomue-verifier
Open-source local verification, published on npm and exercised across three operating systems and two Node.js versions.
Machine-readable evidence
Scoped results, versions, reasons, and boundaries can travel with an agent workflow instead of becoming uncheckable prose.
Available now — release candidate
Agents can call nomue Record Verifier over MCP
The exact @licklider/nomue-verifier-mcp@0.2.0-rc.1 release candidate is public on npm and in the official MCP Registry as nomue Record Verifier. It exposes the method-neutral verify_nomue_record tool over local stdio; the current supported scientific scope remains Release 1 Welch Record verification. It delegates to @licklider/nomue-verifier@0.2.1-rc.1 and has passing package-path CI across Linux, macOS, and Windows. It requires no account, API key, environment variable, or Licklider-hosted service. The first npx launch may download npm dependencies; after installation, verification runs locally. This release candidate supports stdio only: it is not a hosted HTTP endpoint and does not add paired-t, Wilcoxon, Mann–Whitney, method selection, raw-sample calculation, or an overall scientific verdict.
Limited release, candidates, and public proposals
Use managed Welch verification and shape the next capabilities
Managed agent-facing Welch workflow — limited Release 1
Approved recipients can submit data and required scientific declarations for a supported Welch calculation, or submit a claimed result with structured evidence for checking, through authenticated MCP and HTTP. The service returns scoped outcomes, reasons, evidence, and next actions.
Use the npm-published Release 1 verifier, the local stdio MCP server, the Protocol, and their machine-readable documentation today. The agent-facing Welch capability is now available in limited Release 1 to approved recipients through authenticated MCP and HTTP interfaces; public self-registration is not available. The hosted capability does not yet emit public Records for replay through the local verifier.
Paired-t / Protocol Release 2
The Release 2 paired-t candidate now has an independently reviewed formal decision-readiness packet. It assembles the D1–D6 decision ledger, numerical and execution evidence, structural candidates, review dispositions, Release 1 safeguards, and the required coupled landing order. The Steward decisions, authoritative issuance, support activation, and release remain open.
Independent groups and multiple comparisons / Protocol Release 3
Release 3 public discussion is open on independent groups and multiple comparisons. The proposal makes design, comparison families, result meaning, and error-control questions explicit across 49 catalogued procedures. Its evidence scope is limited to supplied originals; method adoption and numerical support remain separate decisions.
The minimum 30-day discussion began September 9, 2026. The earliest decision is October 9 at 11:50:18 UTC. Catalogue entries remain review questions; the window does not promise adoption or support.
Two-factor experiments / Protocol Release 4
Release 4 public discussion is open for a balanced two-by-two fixed-factor proposal. Its unissued numerical, report and controlled-execution candidate has reached independently reviewed final readiness, without establishing Protocol support. A separate amendment discussion covers strict binary64 comparison and completed indeterminate results; it changes neither the current verifier nor the original RFC clock.
The minimum 30-day discussion began September 9, 2026. The earliest decision is October 9 at 05:59:47 UTC; the window does not promise adoption or a product release.
The D01/D07 amendment has its own 30-day clock, controlled by GitHub's September 18 creation time of 01:51:08 UTC; the earliest unified decision is October 18 at 01:51:08 UTC. This corrects the earlier body-declared time by nineteen seconds without resetting the original window. The formal decision packet has completed its review; adoption and support remain separate.
Scope and evidence · Public discussion · Numerical scaling study
Shared design and timing evidence / Protocol Release 5
Release 5 public discussion is open on a common evidence view for declared study design and selection timing across analysis families. The proposal covers versioned mappings, timing declarations, explicit limits on what a passing check means, and a shared report view. All three candidate families require separately accepted successors; no new verification capability is available.
The initial scope considers independent two-group, paired two-condition and independent multi-group continuous outcomes. Each depends on its own accepted Contract, Profile, schema and bundle; the factorial proposal remains separate.
The minimum 30-day discussion began September 17, 2026 at 00:38:15 UTC. The earliest decision is October 17 at 00:38:15 UTC, subject to the final impact assessment and remaining holds.
Controlled agent evaluation
The evaluation program examines nomue through decision quality, cost and time, and concrete cases, with explicit comparisons and linked reproduction materials. Current evidence includes two developer-led preprints on constructed Welch workflows and selected cases from the whole-submission study. Results remain specific to their tasks, models and configurations; public archives support reproducing disclosed results, not independently rerunning the private nomue implementations.
Planned platform evolution
Make each verification call portable, persistent, and reusable
Record assembly and emission has entered development. The adopted initial invite-free release plan targets Welch, independent multi-group and paired two-condition capabilities; delivery and scientific activation remain future gates, with no promised date.
01
Independently verifiable Records
Carry exact capability, engine, Protocol, and result evidence into a Record that another party can check without trusting the product runtime.
02
Agent-native Project state
Let an authorized agent return to the same research task without rebuilding canonical context from chat history.
03
Persistent resumability
Reconnect and continue long-running research work while keeping project truth separate from conversational memory.
04
Generalized capability kernel
Stabilize the common contract, result, refusal, evidence, and version surfaces after materially different methods prove what is shared.
Planned scientific expansion
Ten method families are already part of the adopted capability architecture
The product plan extends beyond independent two-group Welch into the following research structures. Each family still earns its own scope, evidence, and release.
- Independent multi-group
- Paired two-group
- Repeated measures
- Factorial and interaction
- Nonlinear and dose response
- Nonparametric rank-based
- Categorical outcomes
- Correlation and linear models
- Survival time-to-event
- Count outcomes
Multi-group preparation now includes reviewed, accepted source characterizations of unequal-variance intervals, control and best-treatment comparisons, and closed-testing graphs. These results clarify the assumptions and outputs a future rule must preserve; they do not select or release those methods.Unequal-variance comparisons ·Control or best? ·Testing graphs.
Longer-term compounding system
Real boundary cases become governed improvements
As real use grows, the planned governed learning loop turns privacy-minimized failures and boundary cases into reviewed evidence, regression cases, and improved future capability versions without treating model self-judgment as scientific ground truth.
This is the platform loop: agents use versioned capabilities, results preserve evidence, failures reveal where the boundary needs work, and reviewed improvements return through new versions.