This work is inspired by, and builds directly on, Individual Parameters in Weight-Sparse Transformers Appear Interpretable by Arnau Marin-Llobet and Stefan Heimersheim (NeurIPS 2026; project site on Weightpedia). It reproduces the paper's metric on the model the paper studies, csp_yolo1 from Gao et al. 2025, Weight-sparse transformers have interpretable circuits (openai/circuit_sparsity). The method, the test and the question are theirs; the experiments below are a measurement made with their test, and we are grateful to the authors for the paper and for Weightpedia.
results.json. Select any figure to inspect it at full size.Repository · Interactive atlas of every weight · Paper (arXiv:2607.02964) · Weightpedia
The question
The paper asks an LLM to write 100 candidate Python rules f(tokens, pos) -> bool for a single weight, then tests whether a rule explains when that one weight matters. For the Gao et al. model its Table 1 reports 15.0% ± 1.7 of weights as robustly interpretable.
We ask a narrower question: can a plain program search do the same, and what happens when the test is run on every weight? There is no LLM in the loop anywhere in this work.
What we did
- Ablate every weight exactly. Zero one weight, replay the model, and record the exact per-token change in cross-entropy (and KL). Replay is bit-identical to a full forward pass with the weight zeroed. All 393,216 MLP weights were profiled this way, on 12 repository-exclusive slices of 16,384 tokens from CodeParrot-clean-valid.
- Search for a rule. A Rust search over a causal grammar (token identity, character classes, windows of up to ±8 tokens) proposes rules for each weight at about 600 weights per second, choosing on a train/validation split of the discovery slice.
- Score with the paper's metric, exactly. Restore the weight only at positions where the rule fires, re-run the rest of the network, and measure cross-entropy on the discovery slice and on ten unseen slices.
The metric is the paper's. With the intact model, the loss increase when the weight is zeroed, the loss when the weight is restored only where holds, and the loss when it is restored only where does not hold:
A rule is interpretable when and its coverage (the fraction of positions where it fires) is in . It is robustly interpretable when this holds across the ten held-out slices. The paper does not spell out its robust null, so we read its Fig. 5 cone as and , and we always report two stricter nulls beside it. Each compares against , the score of a random position mask with the same per-slice coverage as the rule:
- Paper's test: and .
- Beats random, same coverage: additionally .
- Strict null: .
Results
| Set of weights | n | Paper's test | Beats random (same coverage) | Strict null |
|---|---|---|---|---|
| Uniform sample | 1,024 | 17.3% ± 1.2 (177) | 10.5% ± 1.0 (108) | 1.8% ± 0.4 (18) |
| High-impact, |ΔCE| ≥ 1e-3 | 1,669 | 40.5% ± 1.2 (676) | 39.2% ± 1.2 (655) | 33.7% ± 1.2 (562) |
| Paper, Table 1 (LLM-written rules) | 15.0% ± 1.7 | not reported | not reported |
Counts of passing weights are in brackets; ± is one binomial SEM. The uniform sample counts every verdict, including weights whose ablation is below the numerical noise floor (17,587 of the 393,216 move the loss by less than 1e-6).
- No LLM was needed for the headline. The search reaches 17.3% ± 1.2 on a uniform sample of 1,024 weights under the paper's test, against the paper's 15.0% ± 1.7.
- That test can be met by broad rules. 65% of the passing rules fire on more than 30% of all tokens, for example "a recent token is alphabetic". Requiring a rule to beat a random mask of the same coverage leaves 10.5%; the strict null leaves 1.8%.
- Where the ablation matters, the rules are narrow and the result holds. For all 1,669 weights whose removal moves the loss by at least 1e-3 nats per token, 40.5% pass, 33.7% pass even the strict null, and the passing rules have a median coverage of 7%. This set is selected by impact, so it is not comparable to the paper's uniform sample.
- Impact helps within the uniform sample too. Of its 104 weights with |ΔCE| ≥ 1e-4, 23.1% pass the paper's test and 7.7% pass the strict null; for the 51 below 1e-6, none do.
What the broad-rule finding means for the test
We read this as a measurement about the test, not a verdict on the paper. The test allows rules that cover up to half of all positions, and in our search most of the rules that pass it are of that broad kind: 65% fire on more than 30% of tokens. A broad rule has a good chance of restoring much of a weight's effect simply because it covers many positions, which is what the coverage-matched random mask checks. Measured that way, the share of passing weights drops from 17.3% to 10.5%, and to 1.8% under the strict null. Where a weight matters a lot, the passing rules are narrow and the drop is small (40.5% to 39.2% to 33.7%).
Three limits on what this shows. We ran a program search, not the paper's LLM; we did not evaluate the paper's rules, so we cannot say how their coverage is distributed. The paper's 15.0% is its own pipeline's number, and our 17.3% uses our reading of its robust test. And our rule selection prices coverage (see the deviations below), which the paper's argmax does not. The comparison is therefore between two methods under one test, plus the same test under tighter nulls, not a replication of the paper's rule set.
A case that holds up: neuron 1863
The paper discusses neuron 1863 as a digit detector. Our search was not told this. For 8 of the neuron's 23 scored weights (7 input, 1 output) it independently finds the rule tokens[pos].isdigit(), which fires on 2.89% of tokens. Their held-out mean score is 0.95 to 1.02 across the ten unseen slices, and all eight pass the strict null.
The other 15 scored weights do not pass. Nine of them also got a digit-based rule (four of those a window such as "a digit appears among the preceding tokens"), but their held-out scores vary widely from slice to slice or do not beat the null; the other six got punctuation, indentation or combined rules. The figure shows all 23 rather than only the eight that worked.
The atlas
The same data is browsable: open the atlas explorer. It covers every one of the 393,216 MLP weights by layer, with each weight's exact ablation effect. Weights that were scored exactly show their rule as executable Python with its discovery and held-out scores and its verdict under all three nulls. Weights without an exact score show the search's best rule, labelled unverified, never as a pass or a failure.
Compute and reproducibility
Everything ran on one Quadro RTX 4000 (8 GB) in fp32: profiling all weights took 50.8 h (with the --fast suffix-MLP path, gated against eager execution's own numerical noise floor), scoring the uniform sample 11.8 h, and scoring the high-impact set 19.4 h, 82.0 h in total. fp16 is not usable: it moves cross-entropy by up to 0.86. Every long stage checkpoints per shard and resumes. All numbers on this page are in assets/results.json, and the formats, metric definitions and validation gates are in CONTRACT.md.
git clone https://github.com/thebasedcapital/sparse-weight-atlas.git
cd sparse-weight-atlas
# first create the .venv and download csp_yolo1, as in the repository README
(cd synth && cargo build --release)
export PYTHONPATH=engine
# the paper's metric on a uniform sample of 1,024 weights
.venv/bin/python -m swa.pilot --model csp_yolo1 --n 1024
# exact scores for every weight with |ΔCE| >= 1e-3
scripts/run_highimpact.sh
The repository README has the full setup, from the environment and model download to the profiling and export stages.
Deviations from the paper
- Rules come from enumerative search over a fixed causal grammar, not an LLM. Candidates are chosen on a train/validation split of the discovery slice, so our InterpA (76.0%, against the paper's 37.3%) is inflated and not comparable.
- Among exactly scored candidates, the chosen rule maximises the score minus half the coverage rather than the plain score.
- The paper does not spell out its robust null. We read its Fig. 5 cone as the paper's test above and always report two stricter coverage-matched random-mask nulls next to it.
- Corpus:
codeparrot/codeparrot-clean-valid, one file per repository, at most 512 tokens per document. - Atlas cards for weights without exact scores show the search's best rule, labelled unverified.
- The high-impact set is selected by impact, so its rates are not comparable to the paper's uniform sample.
Credits and licences
- Method and metric: Marin-Llobet and Heimersheim, arXiv:2607.02964 (CC BY 4.0).
- Model, tokenizer and inference code: Gao et al. 2025, openai/circuit_sparsity, vendored under Apache-2.0. Model weights are downloaded, not redistributed.
- Corpus: codeparrot/codeparrot-clean-valid, downloaded at build time, not redistributed.
- Everything else: MIT. Read the implementation and full results or report an issue.
Cite this entry
@misc{tbc_sparse_weight_atlas,
title = {Sparse weight atlas: every MLP weight of a weight-sparse transformer, explained by an executable rule},
author = {{TBC Research}},
year = {2026},
note = {F-011, catalogue version 1},
howpublished = {\url{https://tbcresearch.org/research/sparse-weight-atlas/}}
}