Skip to content
Mutuus Team8 min read

We Raised Our Own Bar: What Counts as Evidence

We wrote down a stricter process for admitting, testing and publishing our own primitives. It may shrink or retire claims in papers we have already written. This post explains the process, what it changes, and what it has not done yet. There are no new benchmark numbers in it.

methodologyresearchprocess
Share
Skim
Read
Deep Dive

We have written down a stricter process for deciding which primitives deserve a paper and which results count as evidence. A candidate now has to be admitted before it is built. Its study is fixed before it is measured. The incumbent gets the same tuning effort as our primitive. The final numbers come from workloads nobody touched during development. No paper has cleared this bar yet, and this post reports no new performance numbers.

Our March post on evaluation described a seven-phase gauntlet with a binary gate at each phase. We still run it. What we found, looking back at our own work, is that the gauntlet starts too late. By phase one we had already picked a candidate, already liked its biology, and already had a rough idea of which benchmark it would win.

That is the weak point. A process that begins after you have fallen for the idea mostly confirms the idea. So we wrote a program (a design document, dated 2026-10-02, still marked as a draft for review) that adds steps before and after the gauntlet, and tightens the middle.

The stated goal is peer review. We want our primitive papers to survive a skeptical reader at a natural computing journal, and those readers have seen plenty of animal-named algorithms. Passing our own gates is not enough. The evidence has to hold up for someone who has no reason to like us.

Why a research blog needs a process post

We publish wins and losses together. That promise is only checkable if the process that produced the numbers is public and strict.

Every benchmark post we have written reports losses beside wins. The Diatom Bitmap post says Roaring wins on point queries and construction. The Nacre Array numbers post does the same. That habit is good, but it does not answer the harder question: how were the comparisons chosen, and who tuned what?

If we tune our own primitive for a week and run the incumbent with its defaults, we can still print an honest table. The table is accurate and the comparison is unfair. The new process is aimed at that gap, not at the table.

Admission comes before the gauntlet

Before a candidate gets a crate, it has to answer seven questions in writing. A human then decides: admit, incubate, or reject.

The seven admission questions are:

  1. Constraint class. What recurring constraint shows up in at least two kinds of systems, and why is it the primary one?
  2. Selection-pressure match. What pressure shaped the biology, what pressure shapes the software, and do they match?
  3. Incumbent failure. Which incumbent, under exactly which condition, struggles, and why does it matter?
  4. Wedge. What is the one decision-changing win? A pile of small wins does not count.
  5. Loss budget. Where will the primitive lose, and why would an adopter still choose it?
  6. Benchmark shape. Which workloads favor it, which favor the incumbent, and which are a negative control where nothing should differ? All three are designed before building.
  7. Kill condition. What measurement would prove the candidate wrong?

Each answer carries evidence, and each piece of evidence is labeled as observed, interpretation, assumption or conviction. A planned test is never allowed to pass as an observed result. The outcome is a short memo with a recommendation and a human decision on record. Incubate means the candidate helps one product or path but its general claim is unproven. Reject means the match is weak, the win is small, or the losses erase the value.

Rejecting is a real outcome. We expect to use it.

Gate 2 in plain terms

Biological analogs only count if they come from living lineages that evolved under the same pressure the software faces. At least two independent lineages are required.

This is the question the metaphor critique is really asking. If a bird and a bat both evolved wings under pressure to stay airborne at low cost, that convergence tells you something about the problem. If your analog is a crystal forming or a magnet aligning, you can borrow a principle from it, but you cannot call it evolution's answer to anything, because nothing selected it.

So each analog in our survey records whether it is living, what pressure it evolved under, and a citation. Non-living analogs can support a principle. They cannot carry Gate 2 on their own. Two cited, living, independent lineages under the matching pressure can.

The method page on this site already asks for convergence across at least two phyla. The new rule is stricter about what counts as an independent lineage.

Study design and a frozen brief

Before the first run, we write down the questions, the workloads, the metric, the evaluator and the statistics. Then we freeze all of it at a git hash.

The brief fixes:

  • Research questions, and for each hypothesis, what would falsify it.
  • Workloads in three families: home field (should favor us), away field (should favor the incumbent), and a negative control.
  • Three partitions of cases: development, selection, and held-out confirmation.
  • The metric and the evaluator, including a version and a digest, so the tuning loop cannot swap in a friendlier yardstick.
  • The statistical protocol: many independent runs, nonparametric tests, effect sizes, and a Holm correction whenever more than one comparison is made.
  • Threats to validity.

Changing a frozen brief is allowed, but only through a revision record, and a revision invalidates every result downstream of it. You cannot quietly adjust the question after seeing the answer. Anything we saw before the freeze is exploratory, and we will call it exploratory.

Tuning parity and held-out confirmation

The incumbent gets its own tuning job with the same budget as our primitive. Final claims come from one run on cases that were never used for tuning.

Two rules do most of the work here.

First, equal effort. If our primitive gets a search over its parameters, so does the baseline, with a matching budget. Otherwise we are measuring tuning effort, not data structures.

Second, a held-out set that is used once. The confirmation cases must be new, they must not have appeared during development or tuning, and each one is recorded as consumed after use. If a primitive wins during tuning and the win shrinks on held-out cases, we report the smaller number. The design document says so directly: wins that shrink under confirmation are acceptable and are reported as measured.

A quick screening margin during tuning is only a screen. It never becomes a claim in a paper.

Red team and reproduction

A manuscript is blocked until a human has attacked it for bias, metaphor, overclaiming and hidden losses, and until a clean checkout reproduces the headline numbers.

Before a manuscript is ready, a person works through a red-team checklist: Does the evaluation measure our thesis or the incumbent's strengths? Is the analogy carrying the argument? Is any claim stronger than its evidence? Does the paper state what it does not claim?

Other publication checks are mechanical. Every reference has to resolve to a real source, and quantitative claims have to be bound to the confirmation snapshot. The AI-use log has to be complete, and the paper's disclosure text is generated from it. A reproduction dry run has to match the headline numbers within a tolerance declared in the brief.

Some checks are hard blocks with no override. Judgment checks, such as asking a model to review for overclaiming, are soft blocks. A model trained on the existing literature carries the incumbent's prior, so we do not let it veto a result on its own.

What it means for our existing papers

Published versions stay as they are. New versions and unreleased drafts have to meet the new bar, and some claims may not survive.

The design document lists four papers.

  • Adaptive Computing Primitives (position paper) is published, and a revision is in progress that adopts the new method.
  • Nacre Array is published on Zenodo. A new version would be the journal manuscript and would face the full bar, including a confirmatory re-run on new held-out workloads.
  • Diatom Bitmap is a draft and has to pass the full bar before release.
  • Mycelial Cache is a draft and has to pass the full bar before release.

Two examples show what this can cost us. For Nacre, the document notes that figures in one internal summary do not match the paper's abstract, and they have to be reconciled before anything else. For Diatom, whether Gate 2 holds depends on which of the paper's analogs are living and evolved under the matching pressure. The design document says magnetic domains and crystal nucleation cannot count. That audit is tracked as an open issue and has not been done. Depending on its outcome, the Diatom paper may be narrowed, redesigned or stopped.

We also found a sentence in our March methodology post that says anyone can clone the repository and reproduce our results. The repository is private, so that is not true today. Whatever we decide about artifact availability, that sentence will be corrected.

What we have not done yet

No paper has cleared the new bar. The second research track has no engine domain yet. This post contains no new benchmark results.

Being plain about status matters more than sounding finished.

  • No paper has been re-released under this process. The process exists as a design, a plan, and an implementation of the Track A states. It is a draft for review.
  • No new benchmark results landed in the last month. The work was process and tooling, so this post has no numbers to report.
  • Track B is undecided. The program also proposes a second track on biological mechanisms for generating novelty in AI, starting with a systematic review. The domain for any engine built from that review is still an open question.
  • Some assumptions are unverified. Examples are the preprint policy of the target journal and whether restricted deposits can be shared with reviewers the way we plan.
  • The public method page has not caught up. It will be updated to describe the admission step once the matching guards are actually enforced, so it never describes a process that is not running.

We would rather publish fewer papers that hold up than more papers that make the same promises the field has already learned to distrust.

The repository behind this work is private for now. When it opens, the design, the guards, and their tests will be there to inspect.

Related Posts

Discussion