Repository navigation
docs(svd_circuits): add demo notebook, slow name-mover parity, and tool docs - #1831
janmenjayap wants to merge 11 commits into
Conversation
Sweeps the top non-degenerate OV directions of gpt2-small L9H9 through the causal gate in keep mode and asserts at least two of them reconstruct the head's behavior better than an arbitrary same-width in-span control. Qualitative by construction: no external reference implementation exists to pin, and the paper publishes no numeric table for this head, so the file reports the observed gate table rather than asserting a threshold. Paired gated/ungated tests pin the gate's polarity in both directions, and a fixed-seed test covers the seed sensitivity of the averaged control. The paper's negative-name-mover sign claim does not reproduce on gpt2-small under this convention, so it is not asserted here.
Adds the second clause of the reconstruction-fidelity check: retaining more OV directions recovers more of the head's output. The recovered quantity is the rank-k reconstruction of the head's OV map, whose relative Frobenius residual is non-increasing in k by Eckart-Young. That is asserted directly, over a ladder that advances block by block so it never splits a degenerate block, which patch_along_directions refuses. The downstream metric is deliberately not asserted to be monotone in k. A metric is a nonlinear functional of the patched logits, so adding a direction can move it either way. Measured on gpt2-small L9H9, the logit difference rises from 0.039 at k=17 to 0.226 at k=29 and peaks at 0.449 at k=33 before falling to zero at k=64, while the residual falls monotonically across the same ladder. Asserting metric monotonicity would encode a claim that is false on a real head, so the test asserts the endpoints instead: one direction moves the metric materially, and retaining the whole span is a no-op.
Adds a runnable walkthrough of the tool plus a reference page, and registers the notebook with the notebook test target and the CI notebook matrix. The demo decomposes gpt2-small layer 9 head 9, prints the rank report, decodes the vocab readout for the top directions, projects the head's real per-position output onto its OV basis, and runs the causal gate in both keep and ablate mode so the polarity flip in `gated` is visible. Its closing section states what the tool does not establish: single head, single model, mechanism rather than the paper's exact subfunction inventory, and seed-sensitive verdicts near the baseline threshold. The docs page covers the QK/OV convention, the degeneracy guard, the mode-dependent gate, and compatibility mode, and links the demo. The analysis tools guide gains a cross-link and now names the causal gate alongside the decomposition, since citing only the decomposition implies the readout is usable on its own. The demo is added to the docs build's notebook copy list so its page renders.
…ells Executes the demo so its outputs are stored, matching every other registered notebook in the repository, and fixes two cells that were wrong. The vocab readout cell indexed `readout[:, i]` as if it held token ids. It does not: `vocab_readout` returns `[d_vocab, k]`, so column i is direction i's projection onto every token and the top tokens are the largest-magnitude entries of that column. The old cell decoded the first ten rows of the projection, which rendered as ten copies of the same token. The two gate cells swept only the first five rank-report rows, which are all degenerate for this head, so the loops printed nothing. They now sweep every attributable direction, which is what makes the gate a result rather than a sample: three directions pass in keep mode and one in ablate mode.
The notebook's stored outputs disagreed with a fresh run on the CI runner in the last printed decimal of three tables: the projection coefficients differed in the third decimal (0.193 vs 0.192) and the gate deltas in the fourth (0.0775 vs 0.0776). Every underlying quantity agrees to two decimals, so this is CPU-kernel float divergence surfacing through the printed precision, not a behavioural difference. Reduces the printed precision rather than widening the shared sanitizer. The sanitizer's rules only truncate floats carrying four or more decimals, so a value printed at three decimals is compared exactly; adding a rule for three would weaken the check for every notebook in the repository to fix one cell. Coefficients now print at two decimals and gate deltas at three. The rank report's sigma column stays at four: it is computed once from the weights, is identical on both platforms, and is the column a reader uses to identify a direction. The gated column is a boolean and is unchanged.
Adding the SVD demo to the notebook-test target joined three recipe lines into
one, so make ran
pytest ... stable_lm.ipynb pytest ... SVD_Circuits_Demo.ipynb pytest ... SVD_Interpreter_Demo.ipynb
as a single command. pytest received the extra paths as positional arguments and
collected all three notebooks anyway, so the target still worked by accident, but
the rerun flags applied once instead of per notebook and the line was unreadable.
Splits it back into one command per line, with the new notebook in alphabetical
position between stable_lm and SVD_Interpreter_Demo. The diff against the base is
now a single added line.
nbval compares stdout exactly, and every cell in this notebook that prints a float32 reduction can differ in its last printed digit across CPU kernels. Reducing the printed precision did not fix it: the projection coefficients then differed in the second decimal (1.75 vs 1.74) and the ablate-mode deltas in the third (-0.016 vs -0.015), because the underlying reductions sit on a rounding boundary and one ulp of difference crosses it. Lowering the precision further would eventually print a constant and stop being informative. Measured margins to the nearest rounding boundary confirm the exposure rather than assuming it. The rank report's sigma/sigma_max sits 8.5e-07 from a boundary at four decimals, the vocab readout's tenth and eleventh tokens are 1.5e-04 apart, and the keep-mode gate's direction 50 is 3.3e-05 from a boundary at three decimals. The keep-mode table had already failed once on direction 55. Marks every float-printing cell NBVAL_IGNORE_OUTPUT, which is the convention the other demos already use for output that is not a pinned value. All seven cells still execute and still print, so the rendered notebook is unchanged for a reader; only the stored-versus-fresh comparison is skipped. nbval still verifies that the notebook runs end to end, and the gate's polarity and count are asserted by the unit, integration, and slow parity tests rather than by this illustration.
|
As is often the case with these PRs, due to size and volume it's going to take me a little while to get through reviewing all 4 thoroughly. I should be able to get reviews in for at least 2/4 of your PRs today, but I will endeavor to get through all 4 if I have time |
jlarson4
left a comment
There was a problem hiding this comment.
Thanks for wrapping up this Proposal! Glad to finally see it come full circle. A couple small comments below
| assert rows, "no attributable OV directions to sweep" | ||
| assert all(math.isfinite(row.delta_metric) for row in rows) | ||
| assert all(math.isfinite(row.baseline_delta_metric) for row in rows) | ||
| assert sum(row.gated for row in rows) >= MIN_GATED_SUBFUNCTIONS |
There was a problem hiding this comment.
This count holds for seed 0 at 16 draws, but not for most other seeds, and not for seed 0 at the tool's default of 32 draws, where the demo gates only direction 49 among these eight. Random in-span directions clear the same bar more often than these eight do, so the count doesn't separate the head's directions from arbitrary ones. Since #1767 framed this check as qualitative, could the test report the counts without asserting them, or assert something random directions fail?
There was a problem hiding this comment.
I've removed both the minimum-count assertion and the separate requirement that at least one direction gate. The sweep now reports the seed, draw count, number of eligible directions, and number gated. It exercises seeds 0, 1, and 2 at both 16 and 32 draws, with 32 as the default. Zero gates is a valid outcome; that case was exercised for seed 1 at both draw counts.
The assertions now check eligibility, finite results, nonnegative control magnitudes, the keep-mode gate comparison, and reproducibility under a fixed seed. There are also deterministic unit checks that equality with the threshold does not pass either mode's strict comparison. The reported counts are descriptive, not evidence that the selected directions are separated from arbitrary in-span controls.
| print(f"{k:3d} {residual:.8f} {delta:.8f}") | ||
|
|
||
| for prev, cur in zip(residuals, residuals[1:]): | ||
| assert cur <= prev + 1e-12 |
There was a problem hiding this comment.
This residual is computed from U, S and V directly, and it falls with k for any SVD, so it never exercises patch_along_directions. The test still passes if the patch keeps the bottom-k directions or a random subspace of the same width. Can the recovery be measured through the patch, so that keeping the wrong directions fails?
There was a problem hiding this comment.
You're right that the old residual only tested SVD truncation. The replacement passes identity-probe head outputs through patch_along_directions and measures the relative Frobenius error of the actual patched output. It checks that output against the top-k map constructed from the original synthetic weight factors, as well as the singular-value tail prediction.
The coverage includes complete degenerate blocks, bottom-k and random-subspace controls, and a no-op control. I also temporarily substituted bottom-k and random projectors in the implementation: both mutations made the recovery tests fail on the patched tensor comparison. Those mutations were restored, and the unmodified tests pass. A separate tiny-Bridge smoke check still covers real hook wiring; it does not assume that downstream logit differences are monotone in k.
| "\n", | ||
| "# readout is [d_vocab, k]: column i holds direction i's projection onto every token,\n", | ||
| "# so the top tokens for a direction are the largest-magnitude entries of that column.\n", | ||
| "for i in range(5):\n", |
There was a problem hiding this comment.
Directions 3 and 4 sit in the degenerate block from the rank report, but this cell and the projection table below still give each one its own column. Direction 3 is the example cell 7 uses for an unattributable claim. Neither vocab_readout nor project_activations refuses degenerate directions, so the refusal wording in cell 7 and in svd_circuits.md line 47 doesn't hold for them. Can these cells stick to attributable directions, as the gate cells already do, and the wording say which tools refuse?
There was a problem hiding this comment.
The vocab readout, activation table, and gate sweeps now share the rank report's isolated, non-null direction selection. The displayed columns retain their original direction IDs rather than relabelling a contiguous prefix. In the executed notebook, the first five displayed directions are 0, 1, 2, 23, and 48, so directions 3 and 4 no longer receive individually interpreted columns. The display cells also assert eligibility and handle an empty selection.
I've made the refusal wording function-specific: logit_signature rejects degenerate/null individual directions; patch_along_directions rejects selections that split a degenerate block and permits complete blocks as subspaces. vocab_readout and project_activations return raw numerical projections and do not enforce that individual-attribution guard. New synthetic tests cover those distinct contracts without changing the public APIs.
| "## 5. What this does and does not show\n", | ||
| "\n", | ||
| "- **Single head, single model.** Everything above is layer 9 head 9 of `gpt2-small`. Nothing here assembles a multi-head circuit.\n", | ||
| "- **Mechanism, not inventory.** The paper's headline subfunction taxonomy is scale- and model-dependent. `gpt2-small` reproduces the *mechanism* (a name-mover splitting into causally distinct directions), not necessarily the paper's exact subfunction inventory at larger scale.\n", |
There was a problem hiding this comment.
On this prompt, zeroing the whole head moves the logit difference by only +0.13 of 3.36, because downstream heads compensate for its +2.1 direct contribution. Every direction's keep-mode delta falls inside the spread of the random controls, so the gate tables don't separate any direction from an arbitrary one. The claim that gpt2-small reproduces a name mover splitting into causally distinct directions goes further than the tables do, and svd_circuits.md lines 125–126 make the same claim in shorter form. Lets adjust the closing section to claim only what the tables show.
There was a problem hiding this comment.
I've removed the mechanism-replication claim and revised the introductions, headings, gate explanations, and closing sections in both the notebook and tool page. The example now claims a single-head decomposition, descriptive projections, and prompt-specific intervention measurements. It explicitly does not establish named subfunctions, statistical separation from arbitrary directions, or replication of the paper's causal taxonomy. Fixing the random seed is described as repeatability, not confirmation.
I also added an executed whole-head context measurement. On this run, the original logit difference is 3.362 and zeroing the head changes it by +0.129. The notebook reports those measurements without interpreting the endpoint's gate flag, since its empty-span control coincides with the intervention. The caveat distinguishes the final-logit effect from a direct write contribution without claiming to identify a particular compensating downstream component. The tables still report mean control magnitudes, not control distributions, so the text does not infer a measured spread or significance result from them.
Description
Third and final PR of the
svd_circuitsvertical slice from #1767. PR1 (#1768) shipped the weight-space core and PR2 (#1775) shipped the readout, projection, and causal gate; this PR adds the validation the issue committed to, plus the demo and reference docs.Four commits:
test(svd_circuits): add slow gpt2-small name-mover oracle-parity checktest(svd_circuits): check top-k recovery converges monotonicallydocs(svd_circuits): add demo notebook and tool docsdocs(svd_circuits): store demo outputs and fix the readout and gate cellsThe demo and docs commits are combined into one, per your suggestion on the issue.
Validation item 2 is now fully delivered
The issue's item 2 had two clauses. Clause 1 (reconstruction to
atol=1e-4) landed in PR1/PR2. Clause 2 (top-k recovery) lands here.One correction to the issue's wording. The issue says keeping top-
kdirections "monotonically increases recovereddelta_metric". That is not what happens, and asserting it would have encoded a false claim. Retaining more directions reconstructs more of the head's true output, so the reconstruction residual falls monotonically, but the downstream metric is a nonlinear functional of the patched logits and can move either way.Measured on
gpt2-smallL9H9, the logit difference rises from 0.039 atk=17to 0.226 atk=29and peaks at 0.449 atk=33before falling to zero atk=64, while the residual falls monotonically across the same ladder:That measurement is what motivated the assertion; it is not the test's own output. The test runs on the no-download
tiny_bridgefixture (rank 16), where the same shape holds at a smaller scale:So the test asserts the residual is non-increasing (Eckart-Young) plus both endpoints: one direction moves the metric materially, and retaining the whole span is a no-op. The metric curve is printed by the test and the finding is recorded in the docstring and commit message rather than hidden.
The ladder advances block by block rather than one direction at a time, because
patch_along_directionsrefuses a retained set that splits a degenerate block, and L9H9's OV SVD has blocks at[3..9],[10..14], and so on.Parity table (commit 1)
Top non-degenerate OV directions of L9H9,
keep=[i]mode,n_baseline=16, seed 0.gatedhere means the direction alone moved the metric less than an arbitrary same-width in-span control.Four gated directions, against a required minimum of two. Paired gated/ungated tests pin the gate's polarity in both directions, so a constant-
Trueor constant-Falsegatedcannot pass.The demo's table differs, and that is the seed sensitivity in action
The demo sweeps every attributable direction at the default
n_baseline=32rather than the test's reduced 16, so its threshold is tighter (0.1272 vs 0.132577) and the gated set shifts:Directions 2, 48, and 50 sit within 0.005 of the threshold and flip between the two settings. This is the documented seed sensitivity, not a discrepancy: the test asserts the count and the polarity, never a specific direction's verdict. Both tables are reported rather than one being picked to look cleaner.
Two bugs found by executing the demo
The notebook initially shipped unexecuted, which hid two real defects. Both are fixed in the fourth commit, and both were invisible to
nbvalbecause with no stored outputs there was nothing to compare against.vocab_readoutreturns[d_vocab, k], so columniis directioni's projection onto every token, and the top tokens are the largest-magnitude entries of that column. The cell decoded the first ten rows of the projection instead, rendering as ten copies of the same token. It now takestorch.topk(readout[:, i].abs(), 10).indices, and direction 2 reads[' Lindsay', ' McKenna', ...]- the name-like signature expected of a name-mover head.continueskipped every one and the loops printed nothing. They now sweep every attributable direction, which is what makes the gate a result rather than a sample.Every registered notebook in the repository stores its outputs; this one now does too.
Negative result: the negative-name-mover sign claim does not reproduce
The plan listed an optional test asserting that a negative name-mover head (L10H7) produces gated directions with
delta_metricof the opposite sign to L9H9's. It does not hold ongpt2-smallunder this convention - both heads produce positive deltas, so there is no sign flip to assert.Per the plan's own instruction, the test was dropped rather than weakened into something that passes vacuously, and the negative result is recorded in the commit message. Reporting it here as required.
Honesty caveats
test_jacobian_lens_oracle_parity.py, there is no external reference implementation to pin: the authors' repo is research-grade and unpinned, and the paper publishes no numeric table for this head. The file reports the observed gate table rather than asserting a threshold.gpt2-small. No multi-head assembly.gpt2-smallreproduces the mechanism (a name-mover splitting into causally distinct directions), not necessarily the paper's exact subfunction inventory at scale.Files
tests/integration/test_svd_circuits_oracle_parity.pytests/unit/tools/test_svd_circuits.pydemos/SVD_Circuits_Demo.ipynbdocs/source/content/svd_circuits.mddocs/source/index.mddocs/source/content/analysis_tools.mdAPI:linemakefile.github/workflows/checks.ymldocs/make_docs.pycopy_demos()allowlisttests/unit/test_make_docs.pyTen files, against the plan's eight. The two extras are
docs/make_docs.py(the demo copy list is a hardcoded allowlist, not a wildcard, so the demo page would not have rendered without it) andtests/unit/test_make_docs.py(which pins that allowlist and failed until updated).Docs note
analysis_tools.mdalready carried the SVD rows and prose from #1782, so this PR adds a dedicated tool page rather than re-adding them, and widens the trailingAPI:line to namepatch_along_directionsalongsidedecompose_head. Citing only the decomposition implies the readout is usable on its own, which is the opposite of the design contract.Type of change
Checklist
Verification
All rows re-run against the final commit on this branch.
pytest tests/unit -qpytest tests/integration/test_svd_circuits.py -qpytest tests/integration/test_svd_circuits_oracle_parity.py -m slow -qpytest --nbval-sanitize-with demos/doc_sanitize.cfg demos/SVD_Circuits_Demo.ipynb(twice)black --check ./isort --check-only ./pycln --checkmypy .build_docs()content/svd_circuits.htmlandgenerated/demos/SVD_Circuits_Demo.htmlemitted, 9 output blocks in the rendered demo, cross-links resolveThe 57 skipped in the unit run are the
slow-marked tests, whichmake unit-testdeselects via-m "not slow"; the pass count is the same either way.Follow-up
Once this merges, #1767 is complete and I will file the Tier 1 / Tier 2 follow-up issue (QK-side causal validation, quantitative separation metric across the full name-mover set, cross-seed stability, multi-head assembly, automated labelling, additional models). The parity table and the
(k, delta_metric)ladder above will be posted there as the baseline that Tier 1 item 2 supersedes.Branch state
Four commits on
docs/svd-circuits-demo, based ondev, ten files changed (+1453 / -3). The demo notebook is executed with outputs stored, matching every other registered notebook in the repository.