tbcresearch.
Menu
← Back to the research index
F-011 / Result

Sparse weight atlas: every MLP weight of a weight-sparse transformer, explained by an executable rule

An exact ablation of all 393,216 MLP weights of OpenAI's csp_yolo1, a Rust rule search with no LLM, and the paper's own held-out test under three nulls. 17.3% of a uniform sample pass the paper's test, 10.5% beat a coverage-matched random rule and 1.8% pass a strict null.

This work is inspired by, and builds directly on, Individual Parameters in Weight-Sparse Transformers Appear Interpretable by Arnau Marin-Llobet and Stefan Heimersheim (NeurIPS 2026; project site on Weightpedia). It reproduces the paper's metric on the model the paper studies, csp_yolo1 from Gao et al. 2025, Weight-sparse transformers have interpretable circuits (openai/circuit_sparsity). The method, the test and the question are theirs; the experiments below are a measurement made with their test, and we are grateful to the authors for the paper and for Weightpedia.

Log-scale bars: 160,570,368 parameters, 807,936 nonzero transformer parameters, 393,216 MLP weights, and how many of those change the loss by at least 1e-6, 1e-5, 1e-4 and 1e-3 nats per token when removed alone: 375,629, 263,834, 40,741 and 1,669. A second bar splits the 807,936 into MLP, attention, embedding and norm parameters. Open full-size figure. Log-scale bars: 160,570,368 parameters, 807,936 nonzero transformer parameters, 393,216 MLP weights, and how many of those change the loss by at least 1e-6, 1e-5, 1e-4 and 1e-3 nats per token when removed alone: 375,629, 263,834, 40,741 and 1,669. A second bar splits the 807,936 into MLP, attention, embedding and norm parameters. Open full-size figure.
The model and what was measured on it. Every one of the 393,216 MLP weights was ablated exactly. Counts come from the project's results.json. Select any figure to inspect it at full size.

Repository · Interactive atlas of every weight · Paper (arXiv:2607.02964) · Weightpedia

The question

The paper asks an LLM to write 100 candidate Python rules f(tokens, pos) -> bool for a single weight, then tests whether a rule explains when that one weight matters. For the Gao et al. model its Table 1 reports 15.0% ± 1.7 of weights as robustly interpretable.

We ask a narrower question: can a plain program search do the same, and what happens when the test is run on every weight? There is no LLM in the loop anywhere in this work.

What we did

  1. Ablate every weight exactly. Zero one weight, replay the model, and record the exact per-token change in cross-entropy (and KL). Replay is bit-identical to a full forward pass with the weight zeroed. All 393,216 MLP weights were profiled this way, on 12 repository-exclusive slices of 16,384 tokens from CodeParrot-clean-valid.
  2. Search for a rule. A Rust search over a causal grammar (token identity, character classes, windows of up to ±8 tokens) proposes rules for each weight at about 600 weights per second, choosing on a train/validation split of the discovery slice.
  3. Score with the paper's metric, exactly. Restore the weight only at positions where the rule fires, re-run the rest of the network, and measure cross-entropy on the discovery slice and on ten unseen slices.

The metric is the paper's. With CE0\mathrm{CE}_0 the intact model, ΔCE\Delta\mathrm{CE} the loss increase when the weight is zeroed, CEf\mathrm{CE}_f the loss when the weight is restored only where ff holds, and CE¬f\mathrm{CE}_{\neg f} the loss when it is restored only where ff does not hold:

rec=1−CEf−CE0ΔCE,inv=1−CE¬f−CE0ΔCE,score=min⁡(rec, 1−inv).\mathrm{rec}=1-\frac{\mathrm{CE}_f-\mathrm{CE}_0}{\Delta\mathrm{CE}},\qquad \mathrm{inv}=1-\frac{\mathrm{CE}_{\neg f}-\mathrm{CE}_0}{\Delta\mathrm{CE}},\qquad \mathrm{score}=\min(\mathrm{rec},\,1-\mathrm{inv}).

A rule is interpretable when score≥0.75\mathrm{score}\ge 0.75 and its coverage (the fraction of positions where it fires) is in (0, 0.5](0,\,0.5]. It is robustly interpretable when this holds across the ten held-out slices. The paper does not spell out its robust null, so we read its Fig. 5 cone as sˉ≥0.75\bar s\ge 0.75 and sˉ−2 SEM(s)>0\bar s-2\,\mathrm{SEM}(s)>0, and we always report two stricter nulls beside it. Each compares against nkn_k, the score of a random position mask with the same per-slice coverage as the rule:

  • Paper's test: sˉ≥0.75\bar s\ge 0.75 and sˉ−2 SEM(s)>0\bar s-2\,\mathrm{SEM}(s)>0.
  • Beats random, same coverage: additionally sˉ−2 SEM(s)>nˉ\bar s-2\,\mathrm{SEM}(s)>\bar n.
  • Strict null: sˉ−2 SEM(s)>nˉ+2 SD(n)\bar s-2\,\mathrm{SEM}(s)>\bar n+2\,\mathrm{SD}(n).

Results

Bar chart of robustly interpretable weights. Uniform sample of 1,024 weights: 17.3% under the paper's test, 10.5% beating a coverage-matched random rule, 1.8% under a strict null; the paper's Table 1 reports 15.0% plus or minus 1.7. High-impact set of 1,669 weights: 40.5%, 39.2% and 33.7%, selected by impact so not comparable to the paper. Open full-size figure. Bar chart of robustly interpretable weights. Uniform sample of 1,024 weights: 17.3% under the paper's test, 10.5% beating a coverage-matched random rule, 1.8% under a strict null; the paper's Table 1 reports 15.0% plus or minus 1.7. High-impact set of 1,669 weights: 40.5%, 39.2% and 33.7%, selected by impact so not comparable to the paper. Open full-size figure.
Same rules, same exact held-out scores, three tests. Error bars are one binomial SEM. The paper's number is drawn only against the uniform sample, the one set it is comparable to.
Set of weights n Paper's test Beats random (same coverage) Strict null
Uniform sample 1,024 17.3% ± 1.2 (177) 10.5% ± 1.0 (108) 1.8% ± 0.4 (18)
High-impact, |ΔCE| ≥ 1e-3 1,669 40.5% ± 1.2 (676) 39.2% ± 1.2 (655) 33.7% ± 1.2 (562)
Paper, Table 1 (LLM-written rules) 15.0% ± 1.7 not reported not reported

Counts of passing weights are in brackets; ± is one binomial SEM. The uniform sample counts every verdict, including weights whose ablation is below the numerical noise floor (17,587 of the 393,216 move the loss by less than 1e-6).

  • No LLM was needed for the headline. The search reaches 17.3% ± 1.2 on a uniform sample of 1,024 weights under the paper's test, against the paper's 15.0% ± 1.7.
  • That test can be met by broad rules. 65% of the passing rules fire on more than 30% of all tokens, for example "a recent token is alphabetic". Requiring a rule to beat a random mask of the same coverage leaves 10.5%; the strict null leaves 1.8%.
  • Where the ablation matters, the rules are narrow and the result holds. For all 1,669 weights whose removal moves the loss by at least 1e-3 nats per token, 40.5% pass, 33.7% pass even the strict null, and the passing rules have a median coverage of 7%. This set is selected by impact, so it is not comparable to the paper's uniform sample.
  • Impact helps within the uniform sample too. Of its 104 weights with |ΔCE| ≥ 1e-4, 23.1% pass the paper's test and 7.7% pass the strict null; for the 51 below 1e-6, none do.

What the broad-rule finding means for the test

We read this as a measurement about the test, not a verdict on the paper. The test allows rules that cover up to half of all positions, and in our search most of the rules that pass it are of that broad kind: 65% fire on more than 30% of tokens. A broad rule has a good chance of restoring much of a weight's effect simply because it covers many positions, which is what the coverage-matched random mask checks. Measured that way, the share of passing weights drops from 17.3% to 10.5%, and to 1.8% under the strict null. Where a weight matters a lot, the passing rules are narrow and the drop is small (40.5% to 39.2% to 33.7%).

Three limits on what this shows. We ran a program search, not the paper's LLM; we did not evaluate the paper's rules, so we cannot say how their coverage is distributed. The paper's 15.0% is its own pipeline's number, and our 17.3% uses our reading of its robust test. And our rule selection prices coverage (see the deviations below), which the paper's argmax does not. The comparison is therefore between two methods under one test, plus the same test under tighter nulls, not a replication of the paper's rule set.

A case that holds up: neuron 1863

Layer-0 neuron 1863 with its 11 scored input weights on the left and 12 scored output weights on the right. Eight weights, seven input and one output, pass the strict null with held-out mean scores from 0.95 to 1.02. The rule found for all eight is tokens[pos].isdigit(), covering 2.89% of tokens. Open full-size figure. Layer-0 neuron 1863 with its 11 scored input weights on the left and 12 scored output weights on the right. Eight weights, seven input and one output, pass the strict null with held-out mean scores from 0.95 to 1.02. The rule found for all eight is tokens[pos].isdigit(), covering 2.89% of tokens. Open full-size figure.
Layer-0 neuron 1863, the paper's digit-detector example (its Fig. 7). The number in each box is the weight's mean score on the ten held-out slices. Filled boxes pass the strict null.

The paper discusses neuron 1863 as a digit detector. Our search was not told this. For 8 of the neuron's 23 scored weights (7 input, 1 output) it independently finds the rule tokens[pos].isdigit(), which fires on 2.89% of tokens. Their held-out mean score is 0.95 to 1.02 across the ten unseen slices, and all eight pass the strict null.

The other 15 scored weights do not pass. Nine of them also got a digit-based rule (four of those a window such as "a digit appears among the preceding tokens"), but their held-out scores vary widely from slice to slice or do not beat the null; the other six got punctuation, indentation or combined rules. The figure shows all 23 rather than only the eight that worked.

The atlas

The same data is browsable: open the atlas explorer. It covers every one of the 393,216 MLP weights by layer, with each weight's exact ablation effect. Weights that were scored exactly show their rule as executable Python with its discovery and held-out scores and its verdict under all three nulls. Weights without an exact score show the search's best rule, labelled unverified, never as a pass or a failure.

Compute and reproducibility

Everything ran on one Quadro RTX 4000 (8 GB) in fp32: profiling all weights took 50.8 h (with the --fast suffix-MLP path, gated against eager execution's own numerical noise floor), scoring the uniform sample 11.8 h, and scoring the high-impact set 19.4 h, 82.0 h in total. fp16 is not usable: it moves cross-entropy by up to 0.86. Every long stage checkpoints per shard and resumes. All numbers on this page are in assets/results.json, and the formats, metric definitions and validation gates are in CONTRACT.md.

git clone https://github.com/thebasedcapital/sparse-weight-atlas.git
cd sparse-weight-atlas
# first create the .venv and download csp_yolo1, as in the repository README
(cd synth && cargo build --release)
export PYTHONPATH=engine
# the paper's metric on a uniform sample of 1,024 weights
.venv/bin/python -m swa.pilot --model csp_yolo1 --n 1024
# exact scores for every weight with |ΔCE| >= 1e-3
scripts/run_highimpact.sh

The repository README has the full setup, from the environment and model download to the profiling and export stages.

Deviations from the paper

  • Rules come from enumerative search over a fixed causal grammar, not an LLM. Candidates are chosen on a train/validation split of the discovery slice, so our InterpA (76.0%, against the paper's 37.3%) is inflated and not comparable.
  • Among exactly scored candidates, the chosen rule maximises the score minus half the coverage rather than the plain score.
  • The paper does not spell out its robust null. We read its Fig. 5 cone as the paper's test above and always report two stricter coverage-matched random-mask nulls next to it.
  • Corpus: codeparrot/codeparrot-clean-valid, one file per repository, at most 512 tokens per document.
  • Atlas cards for weights without exact scores show the search's best rule, labelled unverified.
  • The high-impact set is selected by impact, so its rates are not comparable to the paper's uniform sample.

Credits and licences

Cite this entry

@misc{tbc_sparse_weight_atlas,
  title = {Sparse weight atlas: every MLP weight of a weight-sparse transformer, explained by an executable rule},
  author = {{TBC Research}},
  year = {2026},
  note = {F-011, catalogue version 1},
  howpublished = {\url{https://tbcresearch.org/research/sparse-weight-atlas/}}
}
← All research