Training transformers with programmatic attention fails without content-dependent programs

I test whether initializing a fresh transformer with the programmatic attention patterns synthesized in Hayes et al.1 confers the convergence speedup that exact attention patterns from a trained GPT-2 do, and if so, what predicts it. Initializing with programs does converge faster than random initialization, but the programs themselves matter little: they speed up training no more than the same programs shuffled onto the wrong documents, and replacing a head with its program changes the trained model's held-out cross-entropy, on average, about as much as deleting the head. The failure appears to be in the programs rather than the interface: a hand-written content-dependent rule, attending just after previous occurrences of the current token, speeds up training +3.9% and collapses to +0.3% when shuffled.

Motivation and commentary

Injecting attention patterns from a trained model into a fresh one accelerates its convergence2. Hayes et al.1 synthesize Python programs that “reproduce those attention maps given only input text”. If the programs capture what the heads do, their patterns should accelerate convergence too.

Call exact the condition where a head is held to the attention the trained model produces on each document, and program the condition where it is held to the best-fitting synthesized pattern. Exact should help; the open question is program. If program beats exact, the fit denoised the learned circuit into a better prior. If it loses most of the benefit, the approximation dropped something that matters for learning.

I trained GPT-2-style models under these initializations and found a large speedup for exact and a modest one for program. I then looked for what predicted the speedup a given head would confer, programmatic or real.

For real heads there are predictors: the benefit rises with depth and with "input dependence", or how much the head's attention moves with the text (ρ = +0.25 over all 144 heads). For programs everything failed. IoU predicted nothing, and sparsity, which did predict program speedup on sink-free heads (ρ = −0.52), turned out to price any sparse pattern rather than these programs: an entropy-matched wrong program did just as well. So the programs' benefit was predictable only from generic properties of the pattern, never from anything tying a program to its head.

Then I tested and falsified the following hypotheses to explain the program's effect on convergence speeds:

hypothesisfalsified by
any stable prior helps causal-uniform over all 144 heads: +0.01%
the IoU selection discards the input-dependence that matters the head's own program, 3.3× the input-dependence of the best-fit program, converges the same
the fixed rule freezes the circuits programs need bias mode reproduces the ordering

Setup

A fresh 12×12×384 GPT-2-style model trains on TinyStories with attention routing supplied per document from cache: one head per run over all 144 GPT-2 heads, and K heads at once up to K = 48, where K = 16 matches the top-16-head patching of Baherwani et al.2 Speedup is steps-to-threshold against an unconstrained baseline, defined only when every threshold is reached.

fixed used for every result here attn = softmax(QKᵀ/√d) attn[h] ← pattern the QK circuit of head h gets no gradient bias a softer variant, empirically ~ fixed logits[h] += λ•log(pattern) attn = softmax(logits) λ is learned, initialised at 1 an all-zero pattern row means the program did not cover that position; the head keeps its own attention at those rows, as in the original repository
The two injection rules. Every result in this report uses fixed.
injected patternmean speedup, one head replaced
(average over all 144 heads)
real attention+3.60%
own program+0.42%
best-fit program (the paper's assignment) +0.34%
causal-uniform+0.01%

Main result

Exact attention transfers and boosts training; programs barely do. Re-pairing patterns with the wrong documents collapses exact's speedup and leaves program runs unmoved. Programs match the heads' sparsity (mean row entropy 1.76 against 1.81), so no property of the patterns themselves explains the gap. What matters is pairing the pattern with the document it was computed from.

heads constrained
sixteen heads constrained at once, strict speedup, 4 seeds real attention +21.3% real, mismatched +2.3% best-fit program +1.9% own program +1.7% causal-uniform    -0.9% +0% +10% +20%
Figure 1: Sixteen heads constrained to their patterns at once (strict speedup over an unconstrained baseline, 4 seeds). The oracle compounds to +21.3% and collapses to +2.3% when its patterns are shuffled; program patterns sit near +2% whether shuffled or not, and causal-uniform below zero. Toggle to K = 1 for the per-head distributions over all 144 heads.
aligned document pattern doc 1 doc 2 doc 3 doc 4 doc 5 @shuf document pattern doc 1 doc 2 doc 3 doc 4 doc 5
Figure 2 Left: each document paired with the pattern computed from it. Right: the same five patterns under a fixed permutation. Shuffling hurts exact attention's speedup but not programmatic. 
−4.85points of speedup lost when a head's exact attention is shuffled across documents
0.00points lost when its best-fit program is shuffled the same way
+3.6 ptsexact over program, paired across all 144 heads, p = 5×10−23

Why exact gives speedups while program doesn't

Any stable prior helps

Causal-uniform is a fixed, valid attention prior injected identically. It buys +0.01% over all 144 heads and −0.9% at K = 16, so the programs' small edge must come from what they specifically contain.

The fit discards the input-dependence that matters

The paper re-ranks candidate programs by document-averaged IoU, which favors flat programs: the chosen fits average 0.049 input-dependence against the library's 0.161, 8.9 SD below a random pick. But giving every head the program synthesized from its own attention, with 3.3× the input-dependence of the best fit, converges the same (+0.42% against +0.34%, p = 0.41) and is equally indifferent to shuffling. A more input-dependent program does not recover the missing document-specific routing.

The fixed rule freezes the circuits programs need

Under bias injection the QK circuits keep their gradients and each head learns a weight λ on its prior. The ordering reproduces (exact +3.90%, own program +0.67%, best-fit +0.30%, uniform +0.05% over 144 heads), and heads amplify their program priors (λ = 1.05–1.17) while gaining nothing from them. So the gap cannot be explained by the model ignoring the program prior.

Testing the pruning explanation

The paper flags that its replacement gains “may be akin to a pruning effect.” I find that they are: a program preserves the trained model on average as well as deleting the head, and the head's own mean pattern does slightly better with no program at all.

cost of replacing one head in the trained GPT-2, nats of cross-entropy head's own mean pattern +0.0029 best-fit program +0.0046 head deleted (mean-ablated) +0.0047 causal-uniform pattern +0.0124 head zeroed +0.0133 0
Interchange interventions on the pretrained GPT-2, 256 held-out documents, mean over the 16 selected heads. Program and deletion are statistically indistinguishable on average (paired difference −0.0001 nats); per head the difference runs from −0.019 to +0.017 nats.

Interpolating between mean and exact attention

The ladder's gentlest replacement was the head's own mean pattern, its shape with the content-dependent routing removed. That template is one endpoint of an interpolation: replace each head's pattern with λ·(its attention on the document) + (1−λ)·(its document-mean template) and train as before (16 heads, 3 seeds, one head per run).

strict speedup as the head's own routing is interpolated back in 0 +2% +4% +6% best-fit program +1.35% +1.5% +3.6% +4.8% +5.7% exact +6.1% 0 ¼ ½ ¾ 1 fraction λ of the head's own routing — λ = 0 its document-mean template, λ = 1 its exact attention
Strict convergence speedup as each head's pattern is interpolated from its document-mean template to its exact attention. Dashed line: the best-fit program.

The template alone matches the best-fit program (+1.51% against +1.35%, p = 0.22), so the program-level speedup is available from the head's own shape with no symbolic fit. Speedup then increases monotonically with the amount of real routing, and every point is significantly above the program (each at p ≤ 4×10−5). The relationship is concave: the first quarter of the routing recovers 47% of the program-to-exact gap, and the last quarter is still significant (exact over λ = ¾ by +0.45 points, p = 3×10−4).

A content-dependent positive control

One remaining possibility is that the fixed interface cannot carry any useful symbolic rule at all, and only the trained model's own attention transfers. To test this I wrote one: attend just after each previous occurrence of the current token, the classic induction pattern, computed from token identities alone. Injected through the same machinery (16 heads, 3 seeds, strict speedup), it speeds up training +3.94% and collapses to +0.33% when its patterns are shuffled across documents, faster aligned than shuffled in 16 of 16 heads (p = 1.5×10−5). It beats the best-fit program by 2.6 points, and a variant attending to the occurrences themselves shows the same signature at +1.80% against +0.37%. On the interpolation curve, this rule lands at the equivalent of λ ≈ 0.32 of the head's real routing. This points to the synthesized program library, rather than the symbolic interface, as the bottleneck.

Open

I have not evaluated downstream abilities like QA, and the positive control covers two rules at one scale. The natural next step is a library of content-dependent rules: previous occurrences, matching brackets, sentence boundaries.

References

  1. Hayes, Li & Andreas. Explaining Attention with Program Synthesis. arXiv:2606.19317.
  2. Baherwani, Chen, Qiu, Wilson & Izmailov. Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns. arXiv:2606.25010.