Skip to content

docs(svd_circuits): add demo notebook, slow name-mover parity, and tool docs - #1831

Open
janmenjayap wants to merge 11 commits into
TransformerLensOrg:devfrom
janmenjayap:docs/svd-circuits-demo
Open

janmenjayap wants to merge 11 commits into
TransformerLensOrg:devfrom
janmenjayap:docs/svd-circuits-demo

Conversation

@janmenjayap

@janmenjayap janmenjayap commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Description

Third and final PR of the svd_circuits vertical slice from #1767. PR1 (#1768) shipped the weight-space core and PR2 (#1775) shipped the readout, projection, and causal gate; this PR adds the validation the issue committed to, plus the demo and reference docs.

Four commits:

Commit What
test(svd_circuits): add slow gpt2-small name-mover oracle-parity check Slow qualitative parity sweep on L9H9
test(svd_circuits): check top-k recovery converges monotonically Closes validation item 2's second clause
docs(svd_circuits): add demo notebook and tool docs Demo notebook, reference page, registrations
docs(svd_circuits): store demo outputs and fix the readout and gate cells Executes the demo and fixes two cells

The demo and docs commits are combined into one, per your suggestion on the issue.


Validation item 2 is now fully delivered

The issue's item 2 had two clauses. Clause 1 (reconstruction to atol=1e-4) landed in PR1/PR2. Clause 2 (top-k recovery) lands here.

One correction to the issue's wording. The issue says keeping top-k directions "monotonically increases recovered delta_metric". That is not what happens, and asserting it would have encoded a false claim. Retaining more directions reconstructs more of the head's true output, so the reconstruction residual falls monotonically, but the downstream metric is a nonlinear functional of the patched logits and can move either way.

Measured on gpt2-small L9H9, the logit difference rises from 0.039 at k=17 to 0.226 at k=29 and peaks at 0.449 at k=33 before falling to zero at k=64, while the residual falls monotonically across the same ladder:

k rel. residual abs(delta_metric)
1 0.9867 0.1543
3 0.9611 0.1685
10 0.8740 0.1029
17 0.7872 0.0385
29 0.6383 0.2252
33 0.5875 0.4487
48 0.3792 0.0442
64 0.0000 0.0000

That measurement is what motivated the assertion; it is not the test's own output. The test runs on the no-download tiny_bridge fixture (rank 16), where the same shape holds at a smaller scale:

k rel. residual abs(delta_metric)
1 0.8736 0.0057
2 0.7898 0.0004
4 0.6370 0.0167
8 0.3490 0.0014
16 0.0000 0.0000

So the test asserts the residual is non-increasing (Eckart-Young) plus both endpoints: one direction moves the metric materially, and retaining the whole span is a no-op. The metric curve is printed by the test and the finding is recorded in the docstring and commit message rather than hidden.

The ladder advances block by block rather than one direction at a time, because patch_along_directions refuses a retained set that splits a degenerate block, and L9H9's OV SVD has blocks at [3..9], [10..14], and so on.


Parity table (commit 1)

Top non-degenerate OV directions of L9H9, keep=[i] mode, n_baseline=16, seed 0. gated here means the direction alone moved the metric less than an arbitrary same-width in-span control.

dir sigma delta_metric baseline_delta_metric gated
0 8.4096 0.154328 0.132577 False
1 8.2600 0.140703 0.132577 False
2 8.1019 0.132237 0.132577 True
23 6.8420 0.193184 0.132577 False
48 5.5535 0.130678 0.132577 True
49 5.4845 0.113163 0.132577 True
50 5.3609 0.131467 0.132577 True
51 5.2535 0.162898 0.132577 False

Four gated directions, against a required minimum of two. Paired gated/ungated tests pin the gate's polarity in both directions, so a constant-True or constant-False gated cannot pass.

The demo's table differs, and that is the seed sensitivity in action

The demo sweeps every attributable direction at the default n_baseline=32 rather than the test's reduced 16, so its threshold is tighter (0.1272 vs 0.132577) and the gated set shifts:

dir sigma delta_metric baseline_delta_metric gated
0 8.4096 0.1543 0.1272 False
1 8.2600 0.1407 0.1272 False
2 8.1019 0.1322 0.1272 False
23 6.8420 0.1932 0.1272 False
48 5.5535 0.1307 0.1272 False
49 5.4845 0.1132 0.1272 True
50 5.3609 0.1315 0.1272 False
51 5.2535 0.1629 0.1272 False
55 5.0464 0.0776 0.1272 True
56 4.9738 0.1140 0.1272 True
59 4.7041 0.1307 0.1272 False
60 4.6103 0.1357 0.1272 False
61 4.4351 0.1718 0.1272 False

Directions 2, 48, and 50 sit within 0.005 of the threshold and flip between the two settings. This is the documented seed sensitivity, not a discrepancy: the test asserts the count and the polarity, never a specific direction's verdict. Both tables are reported rather than one being picked to look cleaner.


Two bugs found by executing the demo

The notebook initially shipped unexecuted, which hid two real defects. Both are fixed in the fourth commit, and both were invisible to nbval because with no stored outputs there was nothing to compare against.

  1. The vocab readout cell indexed the wrong axis. vocab_readout returns [d_vocab, k], so column i is direction i's projection onto every token, and the top tokens are the largest-magnitude entries of that column. The cell decoded the first ten rows of the projection instead, rendering as ten copies of the same token. It now takes torch.topk(readout[:, i].abs(), 10).indices, and direction 2 reads [' Lindsay', ' McKenna', ...] - the name-like signature expected of a name-mover head.
  2. Both gate cells swept only the first five rank-report rows, which are all degenerate for L9H9, so the continue skipped every one and the loops printed nothing. They now sweep every attributable direction, which is what makes the gate a result rather than a sample.

Every registered notebook in the repository stores its outputs; this one now does too.


Negative result: the negative-name-mover sign claim does not reproduce

The plan listed an optional test asserting that a negative name-mover head (L10H7) produces gated directions with delta_metric of the opposite sign to L9H9's. It does not hold on gpt2-small under this convention - both heads produce positive deltas, so there is no sign flip to assert.

Per the plan's own instruction, the test was dropped rather than weakened into something that passes vacuously, and the negative result is recorded in the commit message. Reporting it here as required.


Honesty caveats

  • Qualitative, not a pinned oracle. Unlike test_jacobian_lens_oracle_parity.py, there is no external reference implementation to pin: the authors' repo is research-grade and unpinned, and the paper publishes no numeric table for this head. The file reports the observed gate table rather than asserting a threshold.
  • Single head, single model. L9H9 on gpt2-small. No multi-head assembly.
  • Mechanism, not inventory. gpt2-small reproduces the mechanism (a name-mover splitting into causally distinct directions), not necessarily the paper's exact subfunction inventory at scale.
  • Seed sensitivity. Directions near the baseline threshold can gate either way across control draws. Every sweep passes an explicit seed, and a fixed-seed reproducibility test covers it.

Files

File Change
tests/integration/test_svd_circuits_oracle_parity.py new, slow parity sweep
tests/unit/tools/test_svd_circuits.py top-k recovery test
demos/SVD_Circuits_Demo.ipynb new, 17 cells
docs/source/content/svd_circuits.md new reference page
docs/source/index.md toctree entry
docs/source/content/analysis_tools.md cross-link + widened API: line
makefile notebook-test registration
.github/workflows/checks.yml notebook-checks matrix
docs/make_docs.py demo added to copy_demos() allowlist
tests/unit/test_make_docs.py expected allowlist updated

Ten files, against the plan's eight. The two extras are docs/make_docs.py (the demo copy list is a hardcoded allowlist, not a wildcard, so the demo page would not have rendered without it) and tests/unit/test_make_docs.py (which pins that allowlist and failed until updated).


Docs note

analysis_tools.md already carried the SVD rows and prose from #1782, so this PR adds a dedicated tool page rather than re-adding them, and widens the trailing API: line to name patch_along_directions alongside decompose_head. Citing only the decomposition implies the readout is usable on its own, which is the opposite of the design contract.


Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • This change requires a documentation update

Checklist

  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • I have not rewritten tests relating to key interfaces which would affect backward compatibility

Verification

All rows re-run against the final commit on this branch.

Command Result
pytest tests/unit -q 6750 passed, 57 skipped, 6 xfailed
pytest tests/integration/test_svd_circuits.py -q 1 passed
pytest tests/integration/test_svd_circuits_oracle_parity.py -m slow -q 4 passed
pytest --nbval-sanitize-with demos/doc_sanitize.cfg demos/SVD_Circuits_Demo.ipynb (twice) 7 passed both runs
black --check . / isort --check-only . / pycln --check clean
mypy . Success, no issues in 385 source files
build_docs() exit 0; content/svd_circuits.html and generated/demos/SVD_Circuits_Demo.html emitted, 9 output blocks in the rendered demo, cross-links resolve

The 57 skipped in the unit run are the slow-marked tests, which make unit-test deselects via -m "not slow"; the pass count is the same either way.


Follow-up

Once this merges, #1767 is complete and I will file the Tier 1 / Tier 2 follow-up issue (QK-side causal validation, quantitative separation metric across the full name-mover set, cross-seed stability, multi-head assembly, automated labelling, additional models). The parity table and the (k, delta_metric) ladder above will be posted there as the baseline that Tier 1 item 2 supersedes.


Branch state

Four commits on docs/svd-circuits-demo, based on dev, ten files changed (+1453 / -3). The demo notebook is executed with outputs stored, matching every other registered notebook in the repository.


Sweeps the top non-degenerate OV directions of gpt2-small L9H9 through the
causal gate in keep mode and asserts at least two of them reconstruct the
head's behavior better than an arbitrary same-width in-span control.

Qualitative by construction: no external reference implementation exists to
pin, and the paper publishes no numeric table for this head, so the file
reports the observed gate table rather than asserting a threshold. Paired
gated/ungated tests pin the gate's polarity in both directions, and a
fixed-seed test covers the seed sensitivity of the averaged control.

The paper's negative-name-mover sign claim does not reproduce on gpt2-small
under this convention, so it is not asserted here.
Adds the second clause of the reconstruction-fidelity check: retaining more OV
directions recovers more of the head's output.

The recovered quantity is the rank-k reconstruction of the head's OV map, whose
relative Frobenius residual is non-increasing in k by Eckart-Young. That is
asserted directly, over a ladder that advances block by block so it never splits
a degenerate block, which patch_along_directions refuses.

The downstream metric is deliberately not asserted to be monotone in k. A metric
is a nonlinear functional of the patched logits, so adding a direction can move
it either way. Measured on gpt2-small L9H9, the logit difference rises from 0.039
at k=17 to 0.226 at k=29 and peaks at 0.449 at k=33 before falling to zero at
k=64, while the residual falls monotonically across the same ladder. Asserting
metric monotonicity would encode a claim that is false on a real head, so the
test asserts the endpoints instead: one direction moves the metric materially,
and retaining the whole span is a no-op.
Adds a runnable walkthrough of the tool plus a reference page, and registers the
notebook with the notebook test target and the CI notebook matrix.

The demo decomposes gpt2-small layer 9 head 9, prints the rank report, decodes the
vocab readout for the top directions, projects the head's real per-position output
onto its OV basis, and runs the causal gate in both keep and ablate mode so the
polarity flip in `gated` is visible. Its closing section states what the tool does
not establish: single head, single model, mechanism rather than the paper's exact
subfunction inventory, and seed-sensitive verdicts near the baseline threshold.

The docs page covers the QK/OV convention, the degeneracy guard, the mode-dependent
gate, and compatibility mode, and links the demo. The analysis tools guide gains a
cross-link and now names the causal gate alongside the decomposition, since citing
only the decomposition implies the readout is usable on its own.

The demo is added to the docs build's notebook copy list so its page renders.
…ells

Executes the demo so its outputs are stored, matching every other registered
notebook in the repository, and fixes two cells that were wrong.

The vocab readout cell indexed `readout[:, i]` as if it held token ids. It does
not: `vocab_readout` returns `[d_vocab, k]`, so column i is direction i's
projection onto every token and the top tokens are the largest-magnitude entries
of that column. The old cell decoded the first ten rows of the projection, which
rendered as ten copies of the same token.

The two gate cells swept only the first five rank-report rows, which are all
degenerate for this head, so the loops printed nothing. They now sweep every
attributable direction, which is what makes the gate a result rather than a
sample: three directions pass in keep mode and one in ablate mode.
The notebook's stored outputs disagreed with a fresh run on the CI runner in the
last printed decimal of three tables: the projection coefficients differed in the
third decimal (0.193 vs 0.192) and the gate deltas in the fourth (0.0775 vs
0.0776). Every underlying quantity agrees to two decimals, so this is CPU-kernel
float divergence surfacing through the printed precision, not a behavioural
difference.

Reduces the printed precision rather than widening the shared sanitizer. The
sanitizer's rules only truncate floats carrying four or more decimals, so a value
printed at three decimals is compared exactly; adding a rule for three would
weaken the check for every notebook in the repository to fix one cell.

Coefficients now print at two decimals and gate deltas at three. The rank report's
sigma column stays at four: it is computed once from the weights, is identical on
both platforms, and is the column a reader uses to identify a direction. The
gated column is a boolean and is unchanged.
Adding the SVD demo to the notebook-test target joined three recipe lines into
one, so make ran

    pytest ... stable_lm.ipynb    pytest ... SVD_Circuits_Demo.ipynb    pytest ... SVD_Interpreter_Demo.ipynb

as a single command. pytest received the extra paths as positional arguments and
collected all three notebooks anyway, so the target still worked by accident, but
the rerun flags applied once instead of per notebook and the line was unreadable.

Splits it back into one command per line, with the new notebook in alphabetical
position between stable_lm and SVD_Interpreter_Demo. The diff against the base is
now a single added line.
nbval compares stdout exactly, and every cell in this notebook that prints a
float32 reduction can differ in its last printed digit across CPU kernels.
Reducing the printed precision did not fix it: the projection coefficients then
differed in the second decimal (1.75 vs 1.74) and the ablate-mode deltas in the
third (-0.016 vs -0.015), because the underlying reductions sit on a rounding
boundary and one ulp of difference crosses it. Lowering the precision further
would eventually print a constant and stop being informative.

Measured margins to the nearest rounding boundary confirm the exposure rather
than assuming it. The rank report's sigma/sigma_max sits 8.5e-07 from a boundary
at four decimals, the vocab readout's tenth and eleventh tokens are 1.5e-04
apart, and the keep-mode gate's direction 50 is 3.3e-05 from a boundary at three
decimals. The keep-mode table had already failed once on direction 55.

Marks every float-printing cell NBVAL_IGNORE_OUTPUT, which is the convention the
other demos already use for output that is not a pinned value. All seven cells
still execute and still print, so the rendered notebook is unchanged for a
reader; only the stored-versus-fresh comparison is skipped. nbval still verifies
that the notebook runs end to end, and the gate's polarity and count are asserted
by the unit, integration, and slow parity tests rather than by this illustration.
@jlarson4

Copy link
Copy Markdown
Collaborator

As is often the case with these PRs, due to size and volume it's going to take me a little while to get through reviewing all 4 thoroughly. I should be able to get reviews in for at least 2/4 of your PRs today, but I will endeavor to get through all 4 if I have time

@jlarson4 jlarson4 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for wrapping up this Proposal! Glad to finally see it come full circle. A couple small comments below

assert rows, "no attributable OV directions to sweep"
assert all(math.isfinite(row.delta_metric) for row in rows)
assert all(math.isfinite(row.baseline_delta_metric) for row in rows)
assert sum(row.gated for row in rows) >= MIN_GATED_SUBFUNCTIONS

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This count holds for seed 0 at 16 draws, but not for most other seeds, and not for seed 0 at the tool's default of 32 draws, where the demo gates only direction 49 among these eight. Random in-span directions clear the same bar more often than these eight do, so the count doesn't separate the head's directions from arbitrary ones. Since #1767 framed this check as qualitative, could the test report the counts without asserting them, or assert something random directions fail?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've removed both the minimum-count assertion and the separate requirement that at least one direction gate. The sweep now reports the seed, draw count, number of eligible directions, and number gated. It exercises seeds 0, 1, and 2 at both 16 and 32 draws, with 32 as the default. Zero gates is a valid outcome; that case was exercised for seed 1 at both draw counts.

The assertions now check eligibility, finite results, nonnegative control magnitudes, the keep-mode gate comparison, and reproducibility under a fixed seed. There are also deterministic unit checks that equality with the threshold does not pass either mode's strict comparison. The reported counts are descriptive, not evidence that the selected directions are separated from arbitrary in-span controls.

Comment thread tests/unit/tools/test_svd_circuits.py Outdated
print(f"{k:3d} {residual:.8f} {delta:.8f}")

for prev, cur in zip(residuals, residuals[1:]):
assert cur <= prev + 1e-12

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This residual is computed from U, S and V directly, and it falls with k for any SVD, so it never exercises patch_along_directions. The test still passes if the patch keeps the bottom-k directions or a random subspace of the same width. Can the recovery be measured through the patch, so that keeping the wrong directions fails?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right that the old residual only tested SVD truncation. The replacement passes identity-probe head outputs through patch_along_directions and measures the relative Frobenius error of the actual patched output. It checks that output against the top-k map constructed from the original synthetic weight factors, as well as the singular-value tail prediction.

The coverage includes complete degenerate blocks, bottom-k and random-subspace controls, and a no-op control. I also temporarily substituted bottom-k and random projectors in the implementation: both mutations made the recovery tests fail on the patched tensor comparison. Those mutations were restored, and the unmodified tests pass. A separate tiny-Bridge smoke check still covers real hook wiring; it does not assume that downstream logit differences are monotone in k.

Comment thread demos/SVD_Circuits_Demo.ipynb Outdated
"\n",
"# readout is [d_vocab, k]: column i holds direction i's projection onto every token,\n",
"# so the top tokens for a direction are the largest-magnitude entries of that column.\n",
"for i in range(5):\n",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Directions 3 and 4 sit in the degenerate block from the rank report, but this cell and the projection table below still give each one its own column. Direction 3 is the example cell 7 uses for an unattributable claim. Neither vocab_readout nor project_activations refuses degenerate directions, so the refusal wording in cell 7 and in svd_circuits.md line 47 doesn't hold for them. Can these cells stick to attributable directions, as the gate cells already do, and the wording say which tools refuse?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The vocab readout, activation table, and gate sweeps now share the rank report's isolated, non-null direction selection. The displayed columns retain their original direction IDs rather than relabelling a contiguous prefix. In the executed notebook, the first five displayed directions are 0, 1, 2, 23, and 48, so directions 3 and 4 no longer receive individually interpreted columns. The display cells also assert eligibility and handle an empty selection.

I've made the refusal wording function-specific: logit_signature rejects degenerate/null individual directions; patch_along_directions rejects selections that split a degenerate block and permits complete blocks as subspaces. vocab_readout and project_activations return raw numerical projections and do not enforce that individual-attribution guard. New synthetic tests cover those distinct contracts without changing the public APIs.

Comment thread demos/SVD_Circuits_Demo.ipynb Outdated
"## 5. What this does and does not show\n",
"\n",
"- **Single head, single model.** Everything above is layer 9 head 9 of `gpt2-small`. Nothing here assembles a multi-head circuit.\n",
"- **Mechanism, not inventory.** The paper's headline subfunction taxonomy is scale- and model-dependent. `gpt2-small` reproduces the *mechanism* (a name-mover splitting into causally distinct directions), not necessarily the paper's exact subfunction inventory at larger scale.\n",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On this prompt, zeroing the whole head moves the logit difference by only +0.13 of 3.36, because downstream heads compensate for its +2.1 direct contribution. Every direction's keep-mode delta falls inside the spread of the random controls, so the gate tables don't separate any direction from an arbitrary one. The claim that gpt2-small reproduces a name mover splitting into causally distinct directions goes further than the tables do, and svd_circuits.md lines 125–126 make the same claim in shorter form. Lets adjust the closing section to claim only what the tables show.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've removed the mechanism-replication claim and revised the introductions, headings, gate explanations, and closing sections in both the notebook and tool page. The example now claims a single-head decomposition, descriptive projections, and prompt-specific intervention measurements. It explicitly does not establish named subfunctions, statistical separation from arbitrary directions, or replication of the paper's causal taxonomy. Fixing the random seed is described as repeatability, not confirmation.

I also added an executed whole-head context measurement. On this run, the original logit difference is 3.362 and zeroing the head changes it by +0.129. The notebook reports those measurements without interpreting the endpoint's gate flag, since its empty-span control coincides with the intervention. The caveat distinguishes the final-logit effect from a direct write contribution without claiming to identify a particular compensating downstream component. The tables still report mean control magnitudes, not control distributions, so the text does not infer a measured spread or significance result from them.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Proposal] SVD Circuits: singular-vector decomposition of a head's QK/ OV into causally-validated subfunctions

2 participants