Skip to content

Status — done against total#

Read at 2026-09-22, after the 4x5 grid completed at 20/20 and the first hour of Phase J (tests/foreman/phase_j/RECORD.md). Every percentage on this page is a count of named gates, not a judgement of effort. Each table lists its own denominator so a reader can disagree with the denominator rather than with the number.

The north star is conjunctive: three conditions that must all hold. A conjunctive goal is not the average of its legs — it is the weakest leg. Both numbers are given below, and the weakest leg is the honest one.


Headline#

reading value
gates passed, all three conditions 8 of 19 — 42%
the goal as actually stated (conjunctive, weakest leg) 30%
the leg that sets it condition 3, carries the weight of self-attention or JEPA

Condition 1 — understands causality#

2.5 of 5 gates — 50%

# gate state evidence
1 An operator whose entries encode path structure, not pairwise affinity done W_ij = G_ij e^{s_ij} / Z_i^β with G_ij = ∏_{k=j+1..i} m_k e^{iθ_k}, implemented and exact at the (β, qk) corners
2 A bed that demands order, separated from a commuting control done H.1 on a generator handed over as bare integers: operator 0.8620 vs DIAG control 0.2860 at matched 404 params, 5/5 seeds, control saturated within 0.031 of its multiset ceiling 0.3110
3 The non-commutative matrix separated from recurrence alone half 84.9% of the gap needs the matrix, 15.1% is recurrence — measured, but the bed cannot separate the effect from its own generator family
4 A second, independent bed carrying the same property not done —
5 The mechanism named and tested not done both named candidates just failed; see condition 2 gate 7

Condition 2 — beats anything before it#

4 of 9 gates — 44%, with four gates now measured as failures rather than open

# gate state evidence
1 A win exists failed −0.2473 eval NLL over the plain softmax twin, 5 seeds — but that twin carries position only through an absolute table. A zero-parameter RoPE twin recovers 96.3 / 99.7 / 101.4% of the win and an ALiBi twin 108.8 / 106.7 / 108.8%, beating the family by 0.016–0.023 nats at 3 of 3 seeds (R-POS, tests/foreman/phase_j/RECORD_N.md). Against the strongest zero-parameter control there is no win
2 At matched parameters, excess reported not hidden done 783 excess per the same per-block formula on all three domains
3 More than one domain done TinyStories −0.2473, WikiText-103 −0.3285, codeparrot −0.3301
4 Against the strongest baseline lacking the property failed the strongest such baseline is no longer the softmax twin: a Forgetting Transformer gate on the twin (727,704 params, same count as a2) recovers 100.2 / 100.3 / 110.9% of C_win at split seeds 0 / 1 / 2
5 A seed interval on every domain not done 5 seeds / 5 seeds / 1 seed — the third domain is a data point, and its record says so
6 Survives a varied-split protocol done C_win = +0.2470, sd 0.0160, 95% CI [+0.2271, +0.2668] excludes zero, 5/5 split seeds, CRN digests byte-identical across all four arms at every seed
7 The mechanism identified done — and it is prior art the win is a data-dependent forget gate: the FoX twin matches the family's per-token gain in every difficulty quintile to within 0.01. Before that: C_phase −0.0124 (p_BH 0.0177) and C_mass −0.0376 (p_BH 0.0201) — both statistically non-zero and both signed against their own mechanism, so neither contributes any part of the win; C_resid is +0.2969, 120.2% of it
8 Against a strong published baseline, not only its own twin failed FoX (arXiv 2503.02130) minus the family: −0.0004 / −0.0008 / −0.0284 at split seeds 0 / 1 / 2 — a tie, a tie, and a loss, at 131.9 s against 514.8 s per cell
9 Matched on compute, not only on tokens and parameters failed a twin grown to hidden 512 (9,976,320 params) at the same steps and tokens finishes in 226 s and beats the family by 0.1672. Before that: the arms match at 3,538 steps and matched tokens, but wall clock is 514.8s for (f) against 76.6s for (a), n = 5 each, ranges 512.4–519.4 and 75.7–77.6 — a ratio of 6.72×

Phase J hour one settled condition 2's loss leg, and against the family. The effect reproduces and is now explained — by a published forget gate that reaches it at 3.9x less wall clock. A plainly larger twin beats it at under half its clock. Condition 2 can no longer pass through language-model loss; only a capability the forget gate lacks (the DO and JEPA rows of Phase J) can reopen it.

Hour two tried DO, and it did not reopen. Trained with every stream from its own chain and read by a frozen linear probe, the arms score 0.1054 (twin), 0.0954 (FoX) and 0.0986 (family) on 50 intervened worlds, where a three-line Dirichlet plug-in reading the same 128 tokens scores 0.0741. Nobody clears the floor, so condition 2 stays closed by absence, not by a passed kill: "consequence" is not struck. The end-to-end head (R-DO-E2E) scored 0.1057 / 0.1028 / 0.1032 — on the null. Both reads were of arms that never learned the task: 150 steps, next-token loss 1.94–1.97 nats against 1.720 for an online Dirichlet count and 2.079 for a unigram. Ruled DO untested, not dead. Admissibility gate for every future DO row: held-out chain NLL ≤ 1.720 before any read counts. Phase J closes as a kernel phase; the forget gate is the only survivor. An earlier DO row trained on a single chain was ruled void: every arm sat on the no-update null 0.07226 to four decimals.

Correction. Earlier versions of this page scored condition 2 at 6 of 8, then 6 of 9. Gates 1, 2, 3, 4 and 6 were the only ones marked done — five, not six. The headline was 9 of 18 (50%), then 9 of 19 (47%), not 10 of 18 (56%) and 10 of 19 (53%). The count was wrong by one from the first version.

Gate 9 did not exist when this page was first written. An adversarial check of the grid found it: C_win carries an unpriced confound as load-bearing as the parameter confound already carried beside C_mass. The win is token-matched and parameter-matched, and the winning arm burns 6.72x the clock to get it. That is wall clock for an unfused operator against fused SDPA on this implementation, not a measured FLOP count — but unpriced either way, and it is the exact axis condition 3 names. The denominator grew from 8 to 9 because of it, which is why this condition fell from 75% to 67% while nothing about the measurements got worse.

One arithmetic caution carried from the same check: C_resid is distinguishable from C_win is not a second finding. C_resid − C_win ≡ −(C_phase + C_mass) = 0.0500 identically, so it is one statement made twice and is not counted twice here.


Condition 3 — carries the weight of self-attention or JEPA#

1.5 of 5 gates — 30%

# gate state evidence
1 Runs at transformer scale not done 128 hidden / 3 layers / 8 heads / ~725k params
2 Cost competitive with FlashAttention not done no kernel; the carry-collapse theorem retired the kernel claim — a single-token path gate telescopes to a per-key vector, one head dimension in stock SDPA at 1.084×. Now carries measured numbers: 6.72× the twin's wall clock at equal steps (family), and the forget gate alone runs 1.99× an explicit-mask twin and 2.84× is_causal per step (R-K5). This torch build has no flash attention; every "fused" figure here is the memory-efficient kernel, so K5 is deferred to a flash build. Addendum L's SU(2) fold, measured on an implementation that computes the contract's product, runs at 1.03–1.09× forward and 1.08–1.09× forward+backward against that build's fastest SDPA (EFFICIENT_ATTENTION), GPU re-run by the Inspector — a cost with no trained capability behind it yet
3 A prediction or consequence capability self-attention lacks not done the prediction leg answered NO: committor resolution tied by eight bins of piece count taken from its own input
4 The block-summary path proven exact done dense vs bucketed at 8.9e-16 / 3.2e-13 / 2.6e-10 for β = 1 / 0.5 / 0, block 8, 15 closed gates
5 Formal backing half 12 Lean files, 6 sorry remaining

Workstream breakdown#

stream done / total % note
Paired 4×5 grid 20 / 20 cells 100% CRN digests verified from the run records, identical across arms at all five seeds; pre-registered KILL did not fire
Domains at full seed count 2 / 3 67% the third is one seed and is recorded as one
Seeds landed across domains 11 / 15 73% if every domain carried five
Seat producer directories reachable in-tree 22 — no denominator; this is a count, not a fraction
Lean obligations closed 6 of 12 files carry a sorry 50% by file 6 sorry total
Commits pushed this run 45 / 45 100% nothing held locally
Phase I.1 sections written 22 — §22 Rows still out is stale: R3 landed and Kaggle has since run three domains
Abstention decile test ran, then retired NEITHER decile gradient 0.113 → 0.341, top-3 share 40.6% < 50%; control failed, then R-STRAT showed abstention survives stratification (ρ with difficulty 0.0125)
Phase J hour one 17 rows, 0 struck — R-FoX PASS ×3, R-TEMP FAIL, R-COMP PASS, R-STRAT survives; three contract figures did not reproduce
Phase J hour two 10 rows, 0 struck — R-XFER PASS (abstention retires), R-K5 FAIL 1.99×; R-DO-LM, R-DO-IC, R-DO-E2E all void — one chain, then 150 steps; DO untested

What the percentages do not mean#

A gate marked done means one measurement passed with its control and its provenance recorded. It does not mean the question behind the gate is closed. Condition 2 stands at 56% with three gates measured false, which is the uncomfortable and accurate reading: the effect is well measured, now understood, and owned by prior art.

Three counts belong beside the percentages and have no denominator: 6 misattributions found and corrected, 16 pre-registered checks that could not have failed and 3 that could not have passed catalogued, and 5 beds that failed at their own floors. Those are the error rate this project is measuring against itself, and a status page that omits them reads better than the work deserves.