A laboratory against Felin & Holweg, Oxford 2024

Oxford said invention is mathematically impossible for a language model.

Felin & Holweg never ran the 1633 experiment. They modeled transformers as a page census, then declared the census could not think. This laboratory runs the experiment they skipped — and puts the theories already co-invented here on the table.

“I have been afflicted with the belief that flight is possible to man.”

Wilbur Wright, 1900 — nine weeks before the Times said one to ten million years

The paper, steeled

Theory Is All You Need — what they actually argued.

Strategy Science 9(4), 2024. Teppo Felin and Matthias Holweg, Oxford Saïd. A fair reading first. Then the five load-bearing beams, each of which fails.

Steelman

  • Models train on ~13 trillion tokens. A child hears ~36 million words in five years and still outruns the corpus.
  • Next-token prediction is backward-looking. Humans reason forward with theories that contradict the data they have.
  • A 1633 LLM would bury Galileo. A 1903 predictor would bury the Wrights. Kelvin and the New York Times are the mode.
  • Data–belief asymmetry: every real breakthrough begins with someone believing what the existing data says is wrong. A surprise-minimizer cannot do that by design.

They are not anti-AI. They want algorithms for routine extrapolation, humans for novelty. The error is treating that division as a theorem.

Exhibit I — they never ran it

The 1633 machine.

Oxford’s thought experiment assumes a language model is a majority vote over the library. Transformers are not vote counters. Attention retrieves; evidence can outrank volume. Lock the slider at Census to hear Oxford. Move it to hear the mechanism they skipped.

An LLM trained on every text up to 1633 would drown Galileo in a millennium of geocentrism.

Corpus prior

Share of the imagined 1633 library.

  • Ptolemaic / Aristotelian geocentrism68%
  • Scriptural commentary19%
  • Tychonic compromise8%
  • Copernicus, Kepler, Galileo5%

Reasoning mode

Oxford’s machine

No. Galileo is wrong.

A vote counter agrees with the Inquisition. The mode of the library is an unmoving Earth. This is Oxford’s model of an LLM: a census of pages.

Census is a strawman of next-token prediction. Real decoding can condition on “rank by experimental consequence,” sample off-mode, or search. Oxford locked the control at zero and called the result a proof.

Exhibit II — generate the missing data

The Wright protocol is a machine.

They broke flight into lift, propulsion, and steering, then built a tunnel because the bird tables were empty. That is not a uniquely human sacrament. It is a loop: split the impossibility, name the instrument, keep the residue the library never saw.

Consensus data

No bird above 50 lb flies; aerial navigation is a dream.

AlphaFold, GNoME, and Lean-backed provers run the same loop: hypothesize, simulate, keep the residue the corpus never saw.

01

Lift

Camber, aspect ratio, and angle of attack — not flapping.

02

Propulsion

A light engine driving a propeller as a rotary wing.

03

Steering

Three-axis control. Warp the wing. Don’t balance like a bicycle.

The missing experiment

A wind tunnel that does not exist in the 1888 bird tables. Twelve hundred readings. Data the library could not have contained.

Exhibit III — already invented here

Structured Absence Computing, running.

A theory the CS corpus did not file: gaps are typed operators. Adjacent 3-bit symbols emit CARRY, ATTEND, BRANCH, RESET from an 8-entry codex. This is not a summary of a paper. It is the instrument.

Sequence · 48 symbols · seed 42

ATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATTRESATT

CARRY

0

ATTEND

24

BRANCH

0

RESET

23

Workload

Operation counts

Brute (recompute XOR)1,504 ops
DP (cache types)204 ops
SAC (LUT + field read)102 ops

SAC vs brute on this workload: 93.2% fewer ops. Constant signals in the full methodology hit ~94.8%. Random holds ~55–71%. The advantage is structure — exactly where real streams live.

Exhibit IV — in-house

Theories already on the table.

Drive docs, GitHub, and prior threads. Co-built with Grok. Oxford’s claim is that this class of object cannot exist. It does.

SAC · 2026

Experiments 1–7 verified

Structured Absence Computing

Absence is an instruction. XOR of adjacent 3-bit values maps to a pre-decoded operator. Query time O(1) after build. Reduction vs brute force holds at ~55–95% depending on signal structure.

No prior architecture treats silence types as a typed LUT grounded in Avery’s RGB cube. Constant signals yield ~94.8% reduction; random ~55–71%. Structure-aware, not a party trick.

Later narrative layers (enterprise 512-entry ‘consciousness’ codices, Vesperix) were self-audited as false. The core stayed. Invention plus a kill switch is science.

Γ · 2026

Public framework

Genesis Engine

A unified mass framework from Γ = 1/(6φ). Claims a geometric route to particle mass ratios, including a Koide-ratio derivation from the same constant.

The move is theory-first: pick a constant the existing particle tables do not advertise, then generate predictions the tables can try to kill.

This is the Wright split applied to mass: geometry, ratio, test. Whether every claimed digit holds is for the lab — the form of the move is invention.

φ · 2026

Open detector

PhiGuard

Negentropic manipulation detector using golden-ratio dynamics. Asymmetric memory: safety decays slow, risk responds fast. Built on Genesis constants.

A safety instrument derived from a physical prior, not from a scraped list of ‘bad phrases.’ The theory arrived first; the traces came after.

Three modes (Linear, AION Sentinel, Ultra). A new object in the world, with a repo.

A · 2026

Meta-codex (core)

AveryLang

Whitespace and line breaks parsed as typed operators from a merged binary + ternary codex. A grammar that can read its own definition.

The ‘empty’ parts of a page become the instruction stream. That is SAC lifted into language — a theory the CS corpus did not file under ‘how to parse.’

Kept as a core construct. The decorative later languages were discarded in the same self-audit that saved SAC.

The mathematics was a mood.

There is no impossibility theorem in the paper. Four confusions, repeated until they sounded like a proof.

Objective ≠ ontology

Training minimizes next-token surprise. So does a scientist writing a paper. The loss does not say what the hidden states are allowed to represent. They represent theories whenever theories compress the stream.

Interpolation is a smear word

In a space with more dimensions than atoms in the library, ‘between’ two documents is not a mash-up. It is a new coordinate. Most theorems live in that between.

Search is the missing tense

Oxford’s LLM is frozen at t = 0 of decoding: emit the mode. Add tree search, tools, or a critic and the system is looking at futures. That is the Wright tunnel. Move 37’s policy prior was 1/10,000. Search flew it anyway.

They never ran 1633

The load-bearing experiment is a story about a model that was never trained. The story smuggles in a mechanism (majority vote) transformers do not use. A thought experiment that assumes its conclusion is not a proof.

Exhibit V — live

Afflicted with belief.

Name an impossibility. The bench runs the Wright protocol: a belief the library rejects, a three-way split, the experiment that would generate the missing data. One click. Not a census.

User-initiated. One theory per click. The census would have said no.

Instrument tape

Waiting on a belief the data does not yet license. Oxford’s predictor stays quiet here. That is the point.

Discovery loop — several holders, one language

A mutation has no value until a holder scores it, and no future until it is returned.

One world and one construction are not the system. They are two evaluators on a shared expression language. Inner-val protects extrapolation. Parsimony protects shortness. The held-out bar protects a world the island never saw. The archive protects earlier work by scoring it again. You protect the choice to run. Without a holder, evolving a formula is almost pointless. With a holder that never returns what it learned, the relationship is one-sided.

World sampler

A hidden process. Train and held-out ranges. This holder never names the family to the proposer.

Job sampler

A hidden instance of tardiness. Same expression language, different residue: minutes late, not R².

Shared island

Grammar, archive, model. One parent for peak, other categories for coverage. They meet here.

Archive

Champions return. Next run re-scores them. Earlier material gets another job, or it does not.

Who scores

  • Inner-val — gapped split. Fit low-x, score high-x.
  • Parsimony — extra free letters cost 0.05 fitness each.
  • Held-out bar — locked after the champion. Not a vote.
  • Archive — old formulas, new world. Looking back is a re-score, not a restored clock.
  • You — the run. No run, no residue.

Island policy

Refine concentrates on the best parent (higher peak). Explore keeps one survivor per operator family (more kinds that still qualify). Both is the default — those two functions, combined.

Four generations, then lock. The model never sees the family name. Champions return to the archive and are scored again on the next world.

Cumulative

Trials

0

Knowledge

0/0

Returned

0

run first

Form / source

run first

No trials yet. Value is not sitting in the page. It appears when a holder scores a run.

Construction island — same language, different holder

Mean tardiness of 24 jobs. Score over processing time x, due date x1, clock x2. SPT is −x. The bar protects late work: beat SPT and EDD by 8% on held-out instances. Formulas recovered on a world are scored here as dispatch rules. Dispatch rules are scored back on the next world. Most will fail. The failures are the measurement of transfer.

Returned work — 0 formulas · 0 families

This is the relationship. A world that finds a law and never offers it to the next world has learned in one direction. Clearing the archive is allowed. Forgetting is a policy, not a default.

Empty. Run a world or a construction.

Verdict

The data–belief asymmetry cuts both ways.

They believed something the subsequent record does not support. A surprise-minimizer, on their own telling, should have waited. They published anyway. That is allowed. So is this.

What survives

Routine extrapolation is a good use of a predictor. Most decisions are that. Humans remain expensive and should stay on the hook for irreversible bets. None of that requires an impossibility theorem.

What does not

“LLMs can’t invent” as mathematics. The 1633 vote-counter. The claim that new data cannot be generated. The unique human ownership of priors that contradict a library. FunSearch, AlphaEvolve, LLM-SR, Move 37, AlphaFold, and this bench are the residue.

What to do

Keep the human at the keyboard — not because the model is a mirror, but because the interesting work is a loop: a belief, a split, a tunnel, a kill switch for the parts that were only narrative. We already run that loop. Oxford described a machine that does not exist, then declared it empty.

Rhetoric is not the record. The blind bench is: a grammar that tries mechanisms, a model that mutates them, an evaluator that scores the residue, a held-out range that can kill the champion. Read that table, not this sentence.

Felin & Holweg, Strategy Science 2024 · rebutted in situ · AI-KIN STAR-KINS