About

Infrastructure for the verification side of AI research

Licklider builds nomue to check AI-generated claims independently. Public support begins with Welch analyses; life sciences, finance, and chemistry are long-term directions.

Available Today

Public local verification and a hosted capability for approved research agents.

Approved recipients can submit data and required scientific declarations for a supported Welch calculation, or submit a claimed result with structured evidence for checking, through authenticated MCP and HTTP. The service returns scoped outcomes, reasons, evidence, and next actions.

Hosted access is limited to approved recipients, with no public registration or public Record emission yet. Read the limited Release 1 announcement.

Anyone can install the public @licklider/nomue-verifier package from npm and run nomue verify locally to check a conforming Release 1 Record for independent two-group continuous outcomes under the two-sided Welch two-sample t procedure. It recomputes the covered numerical quantities and returns a machine-readable report of the scoped checks without calling a nomue server after installation.

Agents can also call the public local MCP release candidate to verify the same Release 1 Records. Verification runs locally after installation, without a Licklider account or hosted service.

Connect through MCP · Run the verifier · How nomue works

What We Do

Research agents can generate analyses, code, figures, and scientific prose. Licklider works on the separate moment when a workflow needs to check one clearly defined part of that work against evidence.

Our first product, nomue, begins with statistical verification. Its public verifier checks supported Records and returns machine-readable results.

Use the npm-published Release 1 verifier, the local stdio MCP server, the Protocol, and their machine-readable documentation today. The agent-facing Welch capability is now available in limited Release 1 to approved recipients through authenticated MCP and HTTP interfaces; public self-registration is not available. The hosted capability does not yet emit public Records for replay through the local verifier.

Platform

Welch is the first working, publicly checkable vertical slice, not the product boundary. The adopted product sequence expands both the platform beneath each call and the scientific methods available through it.

Licklider is building shared infrastructure for verification calls across AI research. Welch is the first working vertical slice of a broader architecture for portable evidence, persistent agent-native project state, resumability, and expanding scientific capabilities.

Licklider's market is the full set of verification calls that arise across AI research, rather than one research-workflow SaaS category. The long-term infrastructure opportunity is broader than the capabilities available today.

Licklider sits between research agents and the evidence they rely on. The same verification layer can serve different agents and research workflows without replacing the general-purpose interface each researcher chooses.

We expand the platform one evidenced capability at a time, so each new method has an explicit contract and an independently checkable basis.

Product and capability roadmap

Research Progress

Building the evidence for broader verification

Method expansion needs both a precise scientific guarantee and a calculation whose behavior can be checked.

We also measure what verification changes for an agent doing research work. In a paired historical-analysis reuse study, adding nomue reduced API cost by 78.1% and task time by 62.9% on 103 pairs the ordinary-tool comparator also answered correctly. The study uses 48 synthetic Welch tasks and one model; it is a preprint, not peer reviewed. Read the efficiency evaluation.

Release 2 · Reviewed candidate

From numerical bounds to a controlled paired-t execution candidate

The paired-t candidate connects matched observations to a p-value and a 95% confidence interval, with numerical error checks and one controlled runtime. Its formal D1–D6 decision packet has passed repair review.

Independently reviewed Release 2 formal decision-readiness package — not adopted, issued, or supported

Release 3 · Public discussion

nomue Protocol opens Release 3 public discussion for independent groups and multiple comparisons

Release 3 public discussion is open on independent groups and multiple comparisons. The proposal makes design, comparison families, result meaning, and error-control questions explicit across 49 catalogued procedures. Its evidence scope is limited to supplied originals; method adoption and numerical support remain separate decisions.

Public discussion open — supplied-source proposal; method adoption and numerical support remain pending

Unequal-variance comparisons: formulas and guarantees

Release 4 · Public discussion

nomue Protocol opens Release 4 public discussion for two-factor experiments

Release 4 public discussion is open for a balanced two-by-two fixed-factor proposal. Its unissued numerical, report and controlled-execution candidate has reached independently reviewed final readiness, without establishing Protocol support. A separate amendment discussion covers strict binary64 comparison and completed indeterminate results; it changes neither the current verifier nor the original RFC clock.

Public discussion open — reviewed unissued candidate; separate D01/D07 amendment window; no Protocol support

SS, SSE, F, and the limits of power scaling

These results demonstrate how we develop those foundations: inspect original sources, compare calculations with exact references, and repair claims when review finds a gap. The Release 3 independent-group, Release 4 two-factor and Release 5 shared design-evidence proposals are open for public discussion. See the multi-group proposal and invitation to comment, or help shape common design and timing evidence. These research, candidate, and discussion milestones remain separate from the released Welch scope and its distinct public-local and limited-hosted access paths.

Shared Language

Verification is useful only when people and machines can tell what a result means. We publish clear distinctions between execution, clarification, unsupported scope, refusal, failed checks, and properties that were not asserted.

A statistical method name is not a verification contract. A useful guarantee also depends on the comparison family, error criterion, assumptions, sidedness, balance conditions, and exact procedure variant.

The purpose is practical: help a research agent avoid silent assumptions and help a researcher see the exact basis and limit of an AI-assisted result. We publish it as shared working language, not as a claim of ownership over how researchers speak.

Decision vocabulary

How We Work

Precision is part of the product. We investigate very small numerical disagreements, build independent checks, define exact claim boundaries, and keep our diagnosis separate from an upstream maintainer's confirmation.

We separate two questions that are often collapsed: whether the recorded computation ran exactly as stated, and how close its result is to the mathematical target. We review those questions separately. If the evidence cannot establish the required boundary, we improve the evidence, narrow the claim, or keep that result out of public support.

When that work finds a defect in scientific software, we report it back to the project. The contribution matters even when the finding never becomes a nomue capability.

Our upstream work includes 14 upstream reports and 5 matching fixes merged upstream. Reports include issue-tracker and maintainer-email submissions; each entry identifies its status and links to available evidence. This work shows how we investigate scientific software and build the evidence behind our verification infrastructure.

Why “Licklider”

In “Man-Computer Symbiosis” (1960), J. C. R. Licklider described a division of labor in which humans set goals and exercise judgment while computers carry out routine work.

As machines take on more generation and execution, researchers become more important, not less. Their judgment needs clearer evidence and better boundaries. Our name refers to that original vision of complementary roles.

Portrait of J. C. R. Licklider

Founder

Tasuku Kobayashi

Founder and CEO

Tasuku Kobayashi is a serial entrepreneur and the Founder & CEO of Licklider Inc.

After leading a business unit at Recruit, one of Japan's largest tech companies and the parent organization of Indeed and Glassdoor, he launched a veterinary medicine startup and founded an AI enterprise, which he exited via M&A. His career has centered on scaling businesses in heavily regulated domains, where trust is a foundational requirement.

Leveraging that expertise, he founded Licklider to establish the de facto standard for a "Verification Layer"—the key to the safe, compliant, and reliable deployment of generative AI in regulated industries, starting with life sciences research, where precision is non-negotiable. His core hypothesis: a verification layer is what users need before they can operate LLMs with confidence, and it will let researchers use AI to dramatically accelerate their intellectual output.

By building an environment where researchers can work with peace of mind, Licklider aims to unlock the next wave of breakthrough inventions that transform the world—serving as the essential engine behind the scenes that advances global health, safety, and cultural prosperity.

Selected work

Company

Company
Licklider, Inc.
Founder and CEO
Tasuku Kobayashi
Headquarters
N&E BLD. 7F, 1-12-4 Ginza, Chuo-ku, Tokyo, Japan
Founded
July 2024