A research record by Teerth Sharma

resolvent

A JEPA planner ranks its candidates one at a time.
This asks what the whole set says.

How topological structure helps JEPA world models decide: a resolvent over the set of candidate consequences, built on D-JEPA, measured on where it helps and where it does not. The story below is the attention family it grew from. Every result that failed stays on the record.

Scroll the story ↓Read it in full

01 · attention

Attention takes one hop.

The last token reads every earlier token directly, and splits its attention among them. The weights are proportions of one budget.

02 · a chain

A chain takes every hop.

A Markov chain relays along a path, and the weight of a route is the product of its steps. The farther back, the more steps multiply.

03 · the resolvent

One inverse sums every route.

(I − gP)⁻¹ = I + gP + (gP)² + … Causal masking makes it exact: one forward substitution, no iteration, no truncation.

04 · the gate

Close one gate. Every path across it is exactly zero.

Not small. Zero. A decay added to the logit, as in ALiBi or the Forgetting Transformer, cannot reach it: eCᵢ−Cⱼ is never zero. Lean: no_prefix_scan_represents_a_zero_gate.

05 · one head

Three switches. Three known operators.

β, g and qk span softmax attention, the unnormalized kernel and the exact path product. Proved containment: three_corners_containment.

06 · what held

It represents order.

Given S5 as bare integers, the operator scores 0.8620. A control that cannot see order stops at 0.2860, within 0.031 of its own ceiling. 404 parameters each, 5 of 5 seeds.

07 · what died

The language-model win belongs to position.

A zero-parameter ALiBi twin recovers 108.8% of the win. A published forget gate recovers 100.2%. And on chess, prediction reached 1.45% of what the bed offers, tied by a histogram.

08 · the proofs

207 theorems hold it up.

166 theorems and 41 lemmas in Lean 4, 523 dependencies between them. Walk the proof atlas →

09 · now

The candidate set, not the sequence.

The resolvent moved from a causal sequence to the set of futures a JEPA planner chooses between, where D-JEPA's bounded operator already works. It helps where the prediction error is shared by every candidate, and nowhere else. The bound results →