Roadmap

From one proved verification call to shared research infrastructure

Welch is the first working, publicly checkable vertical slice, not the product boundary. The adopted product sequence expands both the platform beneath each call and the scientific methods available through it.

The stages below separate public artifacts, release candidates, next release work, public discussions, and adopted plans. Protocol proposals advance on their own evidence; Release 3, Release 4 and Release 5 discussions are open with distinct scopes and evidence requirements. These stages do not promise product ship dates.

Product development since limited Release 1

Since limited Release 1, nomue development has expanded candidate calculation range, improved completion of difficult calculations, clarified what was checked, and added historical-result handling.

Development update: integrated candidates and completed internal milestones are not a new hosted release or an expansion of public verifier support.

Making verification results more useful to research agents

Available now

A public trust layer for Welch verification

Anyone can install the public @licklider/nomue-verifier package from npm and run nomue verify locally to check a conforming Release 1 Record for independent two-group continuous outcomes under the two-sided Welch two-sample t procedure. It recomputes the covered numerical quantities and returns a machine-readable report of the scoped checks without calling a nomue server after installation.

nomue Protocol Release 1 Public Draft

Public Record, bundle, check, and result meaning for the first supported scope.

Protocol source

@licklider/nomue-verifier

Open-source local verification, published on npm and exercised across three operating systems and two Node.js versions.

npm package

Machine-readable evidence

Scoped results, versions, reasons, and boundaries can travel with an agent workflow instead of becoming uncheckable prose.

Run the verifier

Available now — release candidate

Agents can call nomue Record Verifier over MCP

The exact @licklider/nomue-verifier-mcp@0.2.0-rc.1 release candidate is public on npm and in the official MCP Registry as nomue Record Verifier. It exposes the method-neutral verify_nomue_record tool over local stdio; the current supported scientific scope remains Release 1 Welch Record verification. It delegates to @licklider/nomue-verifier@0.2.1-rc.1 and has passing package-path CI across Linux, macOS, and Windows. It requires no account, API key, environment variable, or Licklider-hosted service. The first npx launch may download npm dependencies; after installation, verification runs locally. This release candidate supports stdio only: it is not a hosted HTTP endpoint and does not add paired-t, Wilcoxon, Mann–Whitney, method selection, raw-sample calculation, or an overall scientific verdict.

0.2.0-rc.1stdioverify_nomue_record3 OS × 2 Node versions

Install and use · npm · Official Registry · Passing CI run

Limited release, candidates, and public proposals

Use managed Welch verification and shape the next capabilities

Managed agent-facing Welch workflow — limited Release 1

Approved recipients can submit data and required scientific declarations for a supported Welch calculation, or submit a claimed result with structured evidence for checking, through authenticated MCP and HTTP. The service returns scoped outcomes, reasons, evidence, and next actions.

Use the npm-published Release 1 verifier, the local stdio MCP server, the Protocol, and their machine-readable documentation today. The agent-facing Welch capability is now available in limited Release 1 to approved recipients through authenticated MCP and HTTP interfaces; public self-registration is not available. The hosted capability does not yet emit public Records for replay through the local verifier.

Read the release announcement

Paired-t / Protocol Release 2

The Release 2 paired-t candidate now has an independently reviewed formal decision-readiness packet. It assembles the D1–D6 decision ledger, numerical and execution evidence, structural candidates, review dispositions, Release 1 safeguards, and the required coupled landing order. The Steward decisions, authoritative issuance, support activation, and release remain open.

Public RFC

Independent groups and multiple comparisons / Protocol Release 3

Release 3 public discussion is open on independent groups and multiple comparisons. The proposal makes design, comparison families, result meaning, and error-control questions explicit across 49 catalogued procedures. Its evidence scope is limited to supplied originals; method adoption and numerical support remain separate decisions.

The minimum 30-day discussion began September 9, 2026. The earliest decision is October 9 at 11:50:18 UTC. Catalogue entries remain review questions; the window does not promise adoption or support.

Scope and evidence · Public discussion

Two-factor experiments / Protocol Release 4

Release 4 public discussion is open for a balanced two-by-two fixed-factor proposal. Its unissued numerical, report and controlled-execution candidate has reached independently reviewed final readiness, without establishing Protocol support. A separate amendment discussion covers strict binary64 comparison and completed indeterminate results; it changes neither the current verifier nor the original RFC clock.

The minimum 30-day discussion began September 9, 2026. The earliest decision is October 9 at 05:59:47 UTC; the window does not promise adoption or a product release.

The D01/D07 amendment has its own 30-day clock, controlled by GitHub's September 18 creation time of 01:51:08 UTC; the earliest unified decision is October 18 at 01:51:08 UTC. This corrects the earlier body-declared time by nineteen seconds without resetting the original window. The formal decision packet has completed its review; adoption and support remain separate.

Scope and evidence · Public discussion · Numerical scaling study

Shared design and timing evidence / Protocol Release 5

Release 5 public discussion is open on a common evidence view for declared study design and selection timing across analysis families. The proposal covers versioned mappings, timing declarations, explicit limits on what a passing check means, and a shared report view. All three candidate families require separately accepted successors; no new verification capability is available.

The initial scope considers independent two-group, paired two-condition and independent multi-group continuous outcomes. Each depends on its own accepted Contract, Profile, schema and bundle; the factorial proposal remains separate.

The minimum 30-day discussion began September 17, 2026 at 00:38:15 UTC. The earliest decision is October 17 at 00:38:15 UTC, subject to the final impact assessment and remaining holds.

Scope and evidence · Public discussion

Controlled agent evaluation

The evaluation program examines nomue through decision quality, cost and time, and concrete cases, with explicit comparisons and linked reproduction materials. Current evidence includes two developer-led preprints on constructed Welch workflows and selected cases from the whole-submission study. Results remain specific to their tasks, models and configurations; public archives support reproducing disclosed results, not independently rerunning the private nomue implementations.

Evaluation design

Planned platform evolution

Make each verification call portable, persistent, and reusable

Record assembly and emission has entered development. The adopted initial invite-free release plan targets Welch, independent multi-group and paired two-condition capabilities; delivery and scientific activation remain future gates, with no promised date.

01

Independently verifiable Records

Carry exact capability, engine, Protocol, and result evidence into a Record that another party can check without trusting the product runtime.

02

Agent-native Project state

Let an authorized agent return to the same research task without rebuilding canonical context from chat history.

03

Persistent resumability

Reconnect and continue long-running research work while keeping project truth separate from conversational memory.

04

Generalized capability kernel

Stabilize the common contract, result, refusal, evidence, and version surfaces after materially different methods prove what is shared.

Planned scientific expansion

Ten method families are already part of the adopted capability architecture

The product plan extends beyond independent two-group Welch into the following research structures. Each family still earns its own scope, evidence, and release.

  • Independent multi-group
  • Paired two-group
  • Repeated measures
  • Factorial and interaction
  • Nonlinear and dose response
  • Nonparametric rank-based
  • Categorical outcomes
  • Correlation and linear models
  • Survival time-to-event
  • Count outcomes

Multi-group preparation now includes reviewed, accepted source characterizations of unequal-variance intervals, control and best-treatment comparisons, and closed-testing graphs. These results clarify the assumptions and outputs a future rule must preserve; they do not select or release those methods.Unequal-variance comparisons ·Control or best? ·Testing graphs.

Longer-term compounding system

Real boundary cases become governed improvements

As real use grows, the planned governed learning loop turns privacy-minimized failures and boundary cases into reviewed evidence, regression cases, and improved future capability versions without treating model self-judgment as scientific ground truth.

This is the platform loop: agents use versioned capabilities, results preserve evidence, failures reveal where the boundary needs work, and reviewed improvements return through new versions.