I test whether initializing a fresh transformer with the programmatic attention patterns synthesized in Hayes et al.1 confers the convergence speedup that exact attention patterns from a trained GPT-2 do, and if so, what predicts it. Initializing with programs does converge faster than random initialization, but the programs themselves matter little: they speed up training no more than the same programs shuffled onto the wrong documents, and replacing a head with its program changes the trained model's held-out cross-entropy, on average, about as much as deleting the head. The failure appears to be in the programs rather than the interface: a hand-written content-dependent rule, attending just after previous occurrences of the current token, speeds up training +3.9% and collapses to +0.3% when shuffled.
Injecting attention patterns from a trained model into a fresh one accelerates its convergence2. Hayes et al.1 synthesize Python programs that “reproduce those attention maps given only input text”. If the programs capture what the heads do, their patterns should accelerate convergence too.
Call exact the condition where a head is held to the attention the trained model produces on each document, and program the condition where it is held to the best-fitting synthesized pattern. Exact should help; the open question is program. If program beats exact, the fit denoised the learned circuit into a better prior. If it loses most of the benefit, the approximation dropped something that matters for learning.
I trained GPT-2-style models under these initializations and found a large speedup for exact and a modest one for program. I then looked for what predicted the speedup a given head would confer, programmatic or real.
For real heads there are predictors: the benefit rises with depth and with "input dependence", or how much the head's attention moves with the text (ρ = +0.25 over all 144 heads). For programs everything failed. IoU predicted nothing, and sparsity, which did predict program speedup on sink-free heads (ρ = −0.52), turned out to price any sparse pattern rather than these programs: an entropy-matched wrong program did just as well. So the programs' benefit was predictable only from generic properties of the pattern, never from anything tying a program to its head.
Then I tested and falsified the following hypotheses to explain the program's effect on convergence speeds:
| hypothesis | falsified by |
|---|---|
| any stable prior helps | causal-uniform over all 144 heads: +0.01% |
| the IoU selection discards the input-dependence that matters | the head's own program, 3.3× the input-dependence of the best-fit program, converges the same |
| the fixed rule freezes the circuits programs need | bias mode reproduces the ordering |
A fresh 12×12×384 GPT-2-style model trains on TinyStories with attention routing supplied per document from cache: one head per run over all 144 GPT-2 heads, and K heads at once up to K = 48, where K = 16 matches the top-16-head patching of Baherwani et al.2 Speedup is steps-to-threshold against an unconstrained baseline, defined only when every threshold is reached.
| injected pattern | mean speedup, one head replaced (average over all 144 heads) |
|---|---|
| real attention | +3.60% |
| own program | +0.42% |
| best-fit program (the paper's assignment) | +0.34% |
| causal-uniform | +0.01% |
Exact attention transfers and boosts training; programs barely do. Re-pairing patterns with the wrong documents collapses exact's speedup and leaves program runs unmoved. Programs match the heads' sparsity (mean row entropy 1.76 against 1.81), so no property of the patterns themselves explains the gap. What matters is pairing the pattern with the document it was computed from.
Causal-uniform is a fixed, valid attention prior injected identically. It buys +0.01% over all 144 heads and −0.9% at K = 16, so the programs' small edge must come from what they specifically contain.
The paper re-ranks candidate programs by document-averaged IoU, which favors flat programs: the chosen fits average 0.049 input-dependence against the library's 0.161, 8.9 SD below a random pick. But giving every head the program synthesized from its own attention, with 3.3× the input-dependence of the best fit, converges the same (+0.42% against +0.34%, p = 0.41) and is equally indifferent to shuffling. A more input-dependent program does not recover the missing document-specific routing.
Under bias injection the QK circuits keep their gradients and each head learns a weight λ on its prior. The ordering reproduces (exact +3.90%, own program +0.67%, best-fit +0.30%, uniform +0.05% over 144 heads), and heads amplify their program priors (λ = 1.05–1.17) while gaining nothing from them. So the gap cannot be explained by the model ignoring the program prior.
The paper flags that its replacement gains “may be akin to a pruning effect.” I find that they are: a program preserves the trained model on average as well as deleting the head, and the head's own mean pattern does slightly better with no program at all.
The ladder's gentlest replacement was the head's own mean pattern, its shape with the content-dependent routing removed. That template is one endpoint of an interpolation: replace each head's pattern with λ·(its attention on the document) + (1−λ)·(its document-mean template) and train as before (16 heads, 3 seeds, one head per run).
The template alone matches the best-fit program (+1.51% against +1.35%, p = 0.22), so the program-level speedup is available from the head's own shape with no symbolic fit. Speedup then increases monotonically with the amount of real routing, and every point is significantly above the program (each at p ≤ 4×10−5). The relationship is concave: the first quarter of the routing recovers 47% of the program-to-exact gap, and the last quarter is still significant (exact over λ = ¾ by +0.45 points, p = 3×10−4).
One remaining possibility is that the fixed interface cannot carry any useful symbolic rule at all, and only the trained model's own attention transfers. To test this I wrote one: attend just after each previous occurrence of the current token, the classic induction pattern, computed from token identities alone. Injected through the same machinery (16 heads, 3 seeds, strict speedup), it speeds up training +3.94% and collapses to +0.33% when its patterns are shuffled across documents, faster aligned than shuffled in 16 of 16 heads (p = 1.5×10−5). It beats the best-fit program by 2.6 points, and a variant attending to the occurrences themselves shows the same signature at +1.80% against +0.37%. On the interpolation curve, this rule lands at the equivalent of λ ≈ 0.32 of the head's real routing. This points to the synthesized program library, rather than the symbolic interface, as the bottleneck.
I have not evaluated downstream abilities like QA, and the positive control covers two rules at one scale. The natural next step is a library of content-dependent rules: previous occurrences, matching brackets, sentence boundaries.