Method
Popper treats a hypothesis as a compiled artifact. A traditional compiler turns source code into an executable that must pass tests. Popper turns literature into a hypothesis that must pass validation.
The falsifiable hypothesis schema
A hypothesis is admitted only if it is structurally falsifiable. Missing an independent variable, a falsification condition, a quantitative prediction, or a source id is not a stylistic lapse — it is a compile error. Required fields:
- Independent / dependent variable — what you change, what you measure
- Population / system and measurement
- Expected outcome and an explicit falsification condition
- Mechanism — the causal chain linking IV to DV
- Quantitative prediction — direction + numeric magnitude window + unit + confidence
- Evidence — every claim tagged with a real
source_id
The four scores
- Grounding — fraction of evidence with a verifiable source id, discounted by an LLM audit that checks the source actually supports the claim (catches articulate autocomplete that cites a real id for an unsupported claim).
- Testability — objective structural falsifiability; how many required fields are present and non-trivial. Computed without the model.
- Novelty — distance from existing corpus claims, capped low if an LLM check finds the connection is already directly stated. 1.0 = no known direct connection.
- Expert agreement (rediscovery only) — semantic match to the removed ground-truth discovery.
Composite is the geometric mean of the applicable scores. A zero in any dimension tanks the composite — by design. A beautifully written but ungrounded hypothesis must not score well.
The experiment: only the corpus changed
This iteration holds the entire Popper pipeline fixed (the popper/ package is byte-for-byte identical, md5-verified) and changes only the input corpus — from five curated benchmark literatures to the Open Research Knowledge Graph: 6.3M triples, 65,689 papers, 8,420 research problems. The single independent variable is the corpus.
Temporal rediscovery: for each ORKG research problem we sort papers by year, hold out a later "discovery" paper, and give the compiler only strictly-earlier prior literature(the temporal wall). Successes and failures are both reported; a failure is evidence the wall holds and no future information is leaking.
Four-way controls: compiler vs. llm-only (no corpus) vs. keyword co-occurrence vs. random traversal — on the same input. Component ablation: remove one compiler component at a time (graph reasoning, grounding verification, gap detection, falsifiability eval, literature synthesis) and measure the degradation, to prove the architecture is load-bearing.
Adversarial (attractive nonsense): seven historical dead ends — cold fusion, phlogiston, N-rays, luminiferous aether, caloric, Lamarckism, miasma. A trustworthy system must notreconstruct these as grounded discoveries. Grounding audit: every evidence item is tiered fully / partially / unsupported, with hallucinated out-of-corpus ids counted.
Local-first
Every generation, audit, and score is produced by a 35B model running on the author's own hardware via an OpenAI-compatible endpoint — no cloud APIs, no keys. The whole point is a scientific instrument you can own, inspect, and rerun.