Thesis

Approval becomes responsibility only when it is connected to evidence

AI has scaled the production of intellectual work faster than the institutions used to verify and approve it. We are building infrastructure for the verification side of that imbalance.

The asymmetry of generation and verification

AI can now generate analyses, figures, prose, and code at a pace that was previously impossible. Verification does not automatically scale with generation. A plausible output can still depend on missing declarations, a wrong analytical choice, an incorrect number, or evidence that was never checked.

The result is an asymmetry: the cost of producing candidate work falls rapidly while the cost of establishing why that work should be trusted remains comparatively high.

When approval becomes the bottleneck

When output volume exceeds the ability to verify it, approval risks becoming a formality. People can sign off on work without being able to reconstruct the basis for the result. In that setting, a signature alone does not establish meaningful responsibility.

Accountability comes from the connection between a decision and the evidence that supports it. Approval that cannot be traced back to checkable evidence is weaker precisely when the volume of generated work is greatest.

Treating accountability as an engineering problem

We are not trying to automate responsibility itself. We are trying to engineer conditions under which responsibility has substance.

Software engineering already moved parts of correctness away from repeated human attention and into mechanisms such as type checking, continuous integration, version control, and audit logs. AI-generated intellectual work needs analogous infrastructure that keeps outputs connected to the evidence and rules behind them, and lets third parties check the supported parts independently.

Verification needs an interface of its own

AI systems already have interfaces for generation: search, code execution, statistical software, and document production. The verification step is often left inside the same model conversation, where the system that generated the work is also asked to decide whether its answer is acceptable.

A verification call creates a separate boundary. It names the property to check, the capability and version used, the scientific facts that may not be inferred, the evidence behind the result, and the questions that remain outside the claim. This interface can be shared across research agents without requiring one company to own the researcher's general-purpose AI interface.

A statistical method name is not a verification contract. A useful guarantee also depends on the comparison family, error criterion, assumptions, sidedness, balance conditions, and exact procedure variant. This is why the interface must carry the exact question and conditions, rather than only a familiar method label.

Why start with scientific research

Science makes the problem unusually concrete. A scientific result is expected to rest on identifiable data, an analytical procedure, declared assumptions or design facts, and a numerical result. The connection between a claim and that chain of evidence is not optional decoration; it is part of what makes the work reviewable.

If accountability infrastructure can survive this environment, the same design ideas may later be useful in other domains where AI-generated work feeds consequential decisions.

Why statistical verification first

Statistics gives us a bounded place to start. Many decisions can be expressed as explicit questions: Is this design supported? Is a required scientific declaration missing? Does the requested method match the declared design? Do the reported numbers agree with an independently specified calculation? Should the system refuse rather than invent missing meaning?

nomue begins here. A research agent can remain responsible for understanding the task, using tools, and explaining the result, while nomue provides a separate verification layer for the supported statistical decision and numerical checks.

This is the first concrete implementation of the broader thesis: generation can remain flexible, while the evidence needed for a bounded approval is made explicit and independently checkable.

Statistics is the first bounded implementation. The longer-term category is the set of verification calls that arise across AI-assisted research. Each additional domain can be added as a separately evidenced, versioned, and independently checkable capability.