Docs

Current capability and boundaries

What can be used now, what comes next, how the platform expands, and how to interpret a successful result.

Available now

Approved recipients can submit data and required scientific declarations for a supported Welch calculation, or submit a claimed result with structured evidence for checking, through authenticated MCP and HTTP. The service returns scoped outcomes, reasons, evidence, and next actions.

Anyone can install the public @licklider/nomue-verifier package from npm and run nomue verify locally to check a conforming Release 1 Record for independent two-group continuous outcomes under the two-sided Welch two-sample t procedure. It recomputes the covered numerical quantities and returns a machine-readable report of the scoped checks without calling a nomue server after installation.

nomue Protocol Release 1 Public Draft and public npm package @licklider/nomue-verifier 0.2.1-rc.1 are the current public artifacts.

Use the npm-published Release 1 verifier, the local stdio MCP server, the Protocol, and their machine-readable documentation today. The agent-facing Welch capability is now available in limited Release 1 to approved recipients through authenticated MCP and HTTP interfaces; public self-registration is not available. The hosted capability does not yet emit public Records for replay through the local verifier.

Public local MCP release candidate

The exact @licklider/nomue-verifier-mcp@0.2.0-rc.1 release candidate is public on npm and in the official MCP Registry as nomue Record Verifier. It exposes the method-neutral verify_nomue_record tool over local stdio; the current supported scientific scope remains Release 1 Welch Record verification. It delegates to @licklider/nomue-verifier@0.2.1-rc.1 and has passing package-path CI across Linux, macOS, and Windows. It requires no account, API key, environment variable, or Licklider-hosted service. The first npx launch may download npm dependencies; after installation, verification runs locally. This release candidate supports stdio only: it is not a hosted HTTP endpoint and does not add paired-t, Wilcoxon, Mann–Whitney, method selection, raw-sample calculation, or an overall scientific verdict.

  • Package: @licklider/nomue-verifier-mcp@0.2.0-rc.1.
  • Start command: npx --yes @licklider/nomue-verifier-mcp@0.2.0-rc.1.
  • Transport: stdio.
  • Tool: verify_nomue_record.
  • Official MCP Registry name: io.github.licklider-ai/nomue-verifier-mcp.
  • Passing package-path CI: Linux, macOS, Windows with Node.js 20 and 22.

Protocol candidates and public discussion

The Release 2 paired-t candidate now has an independently reviewed formal decision-readiness packet. It assembles the D1–D6 decision ledger, numerical and execution evidence, structural candidates, review dispositions, Release 1 safeguards, and the required coupled landing order. The Steward decisions, authoritative issuance, support activation, and release remain open.

Release 3 public discussion is open on independent groups and multiple comparisons. The proposal makes design, comparison families, result meaning, and error-control questions explicit across 49 catalogued procedures. Its evidence scope is limited to supplied originals; method adoption and numerical support remain separate decisions.

Release 4 public discussion is open for a balanced two-by-two fixed-factor proposal. Its unissued numerical, report and controlled-execution candidate has reached independently reviewed final readiness, without establishing Protocol support. A separate amendment discussion covers strict binary64 comparison and completed indeterminate results; it changes neither the current verifier nor the original RFC clock.

Release 5 public discussion is open on a common evidence view for declared study design and selection timing across analysis families. The proposal covers versioned mappings, timing declarations, explicit limits on what a passing check means, and a shared report view. All three candidate families require separately accepted successors; no new verification capability is available.

  • An RFC is a review record, not a support declaration.
  • Implementation evidence does not by itself create a public Protocol capability.
  • A future capability is not an alias or fallback for the exact Release 1 bundle.

How the platform expands

Welch is the first working, publicly checkable vertical slice, not the product boundary. The adopted product sequence expands both the platform beneath each call and the scientific methods available through it.

Licklider is building shared infrastructure for verification calls across AI research. Welch is the first working vertical slice of a broader architecture for portable evidence, persistent agent-native project state, resumability, and expanding scientific capabilities.

Licklider's market is the full set of verification calls that arise across AI research, rather than one research-workflow SaaS category. The long-term infrastructure opportunity is broader than the capabilities available today.

  • Independent multi-group;
  • Paired two-group;
  • Repeated measures;
  • Factorial and interaction;
  • Nonlinear and dose response;
  • Nonparametric rank-based;
  • Categorical outcomes;
  • Correlation and linear models;
  • Survival time-to-event;
  • Count outcomes;

What a successful verification means

A successful result means that the named checks passed under the stated procedure, evidence, and version. The following questions remain outside that bounded result:

  • the truth of input data or researcher declarations;
  • the overall correctness of a research project;
  • the truth of a scientific or causal conclusion;
  • validity for a method outside an explicitly supported scope;
  • publication, regulatory, or clinical acceptance;

Evaluation program

The evaluation program examines nomue through decision quality, cost and time, and concrete cases, with explicit comparisons and linked reproduction materials. Current evidence includes two developer-led preprints on constructed Welch workflows and selected cases from the whole-submission study. Results remain specific to their tasks, models and configurations; public archives support reproducing disclosed results, not independently rerunning the private nomue implementations.

Flat summary for LLMs

This llms.txt-style summary is intentionally repetitive. Use the linked Markdown page when a flat, low-markup representation is preferable.

  • Run the public verifier locally for independent two-group continuous outcomes under the two-sided Welch two-sample t procedure.
  • The Release 1 verifier and local stdio MCP server are public on npm. The agent-facing Welch capability is available in limited Release 1 to approved recipients through authenticated MCP and HTTP interfaces; public registration is not available.
  • Paired t is open Release 2 RFC work, not current Release 1 support.
  • Release 4 public discussion is open for a balanced two-by-two fixed-factor proposal. Its unissued numerical, report and controlled-execution candidate has reached independently reviewed final readiness, without establishing Protocol support. A separate amendment discussion covers strict binary64 comparison and completed indeterminate results; it changes neither the current verifier nor the original RFC clock.
  • Release 5 public discussion is open on a common evidence view for declared study design and selection timing across analysis families. The proposal covers versioned mappings, timing declarations, explicit limits on what a passing check means, and a shared report view. All three candidate families require separately accepted successors; no new verification capability is available.
  • Verification calls across AI research describe the platform category; each additional capability must be evidenced and released separately.
  • A supported verification does not establish source-data truth, overall research correctness, causal truth, or publication acceptance.
  • A controlled comparison is being built before a performance result is published.

Markdown version of this page