Repository navigation
Conversation
…l-seeding fix sparse probe control seeding
…1820) * Add tie-aware roc_auc and average_precision to sparse probe metrics F1 and the other threshold metrics are read at logit >= 0 only, so a probe on a dead (constant) coordinate that predicts every held-out example positive scores F1 = 2p/(1+p). Add threshold-free, tie-aware ROC-AUC (Mann-Whitney with average ranks) and average precision (tied logits form one threshold) to SparseProbeMetrics, propagate both to SparseProbeControl, and document the degenerate case in the sparse probing guide. Fixes #1814 * Clarify the single-class policy for threshold-free probe metrics
…mpute dtype (#1829) * test(projection_kernel): reproduce the half-precision rank collapse The default relative rank tolerance is floored at the storage dtype's epsilon. For bfloat16 storage that floor is about 7.8e-3, which is far above the singular-value ratios that separate a full-rank head from a numerically rank-deficient one. A full-rank head whose spectrum decays gradually therefore collapses to a much smaller measured rank, and a caller requesting the full width fails with "requested rank N exceeds measured rank M". Add a fixture that builds a full-rank matrix with a geometric singular-value spectrum, plus a single-head bridge that injects it as W_Q. Cover the collapse at both half-precision storage dtypes, confirm a structurally rank-deficient half-precision matrix is still detected, and pin the float32 path as unchanged so the tests distinguish half-precision storage from genuinely deficient data. These tests fail against the current default and pass once the tolerance is derived from the compute dtype. * fix(projection_kernel): derive the default rank tolerance from the compute dtype The default relative rank tolerance was floored at the storage dtype's epsilon. For bfloat16 storage that floor is about 7.8e-3, which is far above the singular-value ratios that separate a full-rank head from a numerically rank-deficient one. A full-rank head whose spectrum decays gradually therefore collapsed to a much smaller measured rank, and a caller requesting the full width failed with "requested rank N exceeds measured rank M". The SVD already promotes float16 and bfloat16 inputs to float32, so the measured singular values carry float32 precision regardless of how the checkpoint was stored. Flooring the tolerance at the storage epsilon discarded that precision and made the rank decision depend on the storage dtype rather than on the numerics actually computed. Derive the default from the compute dtype instead, and drop the storage-dtype helper that existed only to feed the old floor. An explicit rtol still takes precedence. One existing deficiency fixture relied on the loose floor to register a near-degenerate direction; it now uses an exactly dependent column so it tests structural deficiency rather than quantization noise. * docs(projection_kernel): describe the compute-dtype default rank tolerance The guide still documented the default relative rank tolerance as `max(max(matrix.shape) * compute_eps, storage_eps)`, which no longer matches the implementation. Rewrite both passages to describe the compute-dtype default and say why it is derived from the dtype the SVD runs in rather than the storage dtype: half-precision inputs are promoted to float32, so a storage-dtype floor would discard the precision of the singular values actually computed and collapse the measured rank of a full-rank half-precision head. Also note that an explicit rtol still takes precedence, and that half-precision checkpoints no longer need a manual override. * test(projection_kernel): cover the half-precision wrapper path on gpt2-small Add an integration case that loads gpt2-small with bfloat16 weights and runs the default O -> Q call without an explicit rtol. gpt2-small is dense MHA with no GQA grouping, so this isolates the storage dtype effect from the head-cardinality effect. The discriminating assertion is the effective tolerance: it must sit below the bfloat16 epsilon, which is what the storage-dtype floor violated. The measured ranks are additionally pinned against the float32 run, and the scores are checked for finiteness and the documented bounds. Exact score equality across dtypes is deliberately not asserted. gpt2-small's worst-conditioned head sits near 40, below the reciprocal of the bfloat16 epsilon, so its measured ranks survive even the storage-dtype floor. The rank and bound checks therefore guard the end-to-end contract rather than reproducing the collapse; a head whose spectrum decays past that reciprocal is covered by the synthetic unit cases.
Co-authored-by: jlarson4 <jonahalarson@comcast.net>
* Forward the requested revision to every remote-code force-import prepare_loading force-imports remote modeling classes so their modules land in sys.modules to patch. Each Hub revision is its own module copy, and only the default revision was imported, so a pinned revision=... loaded by from_pretrained afterwards got an unpatched copy. #1810 fixed this for restore_default_rope_init; do the same for Dream's DreamGenerationConfig and DreamAttention imports and for the BD3LM, GIDD, InternLM2, OpenELM, Ouro, Raven and RWKV-7 adapters. * Forward the requested revision in Baichuan's force-import
* fix(sparse_probing): add std_floor to standardization Near-constant columns were being divided by their own tiny std, blowing noise up to unit scale. The scale now has a floor (default 1e-3, same as the reference). Zero-variance columns still get scale 1, and std_floor=0 keeps the old behavior. Also document the max_eval budget in the max_iter docstrings. * sparse_probing: explain zero-variance scale rule and test setup * sparse_probing: address review feedback Test that sweep_sparse_probe passes std_floor to its main and control fits, fix the zero-variance guard comment, and note that the line search can go one evaluation past the max_eval budget. * sparse_probing: check sweep's default std_floor
* Add Muse Glimmer architecture adapter Bridge MuseGlimmerForConditionalGeneration (text path through model.language_model and lm_head, vision tower and projector as opaque components), register it in the factory, model registry, multimodal list and HF model_type map, and pass output_multiplier through to the bridge config. The unit test compares bridge logits and attention patterns against HF on a tiny random config and skips when transformers lacks muse_glimmer. * Pin the Muse Glimmer output-logit contract The tiny Muse Glimmer model's logits are too small for the tanh cap to change anything, so the forward-parity test cannot catch a missing post-unembedding transform. Add a contract case alongside Granite and Falcon-H1: - nested text_config: final_logit_softcapping and output_multiplier both reach the bridge config; - apply_output_logits_transform applies the multiplier before the softcap (20 * tanh(0.5 * logits / 20)), with logits large enough that dropping the factor fails the comparison. * Expose the merged vision embedding on Muse Glimmer's vision projector HF runs vision_adapter -> vision_projection -> perception_emb_norm and scatters only the last value into the text embeddings, so the vision_projector bridge now wraps perception_emb_norm instead of the raw vision_projection Linear. Its hook_out is then the embedding HF actually merges, matching the Gemma 3 and LLaVA adapters; hook_in is the projection output. Add a structural assertion to the adapter test.
pyproject allows beartype>=0.14.1, but uv.lock still pins 0.14.1. Newer beartype (0.22.x, what a fresh install resolves to) checks dict values and container items, and the jaxtyping test hook then rejects three call sites: - TLWorkerExtension.tl_read_batched_captures returns encode_tensor() dicts, not tensors; annotate it like tl_read_captures. - select_displacement_matched_control_token is called with the decomposition's support tensor; accept a tensor for active_support. - test_resolve_state_dict_key_dense_mlp_fallback passed ints as state-dict values; use tensors.
* Fix recursive torch dispatch for CompositionScores sequences * Consolidate CompositionScores regression tests in protocol class
* wire LLaVA vision hook aliases * fix(vision): preserve CLIP layer norm epsilon
* fix(bridge): preserve seq2seq state in streaming generation * fix(bridge): align seq2seq stream setup and stopping guards
…1845) DynamicCache(config=...) on newer transformers (5.17) compares config.num_kv_shared_layers against an int, which fails on a bare MagicMock config. Use a tiny real NemotronHConfig instead.
* Fix stack_neuron_results for an integer neuron_slice stack_neuron_results wrapped neuron_slice with Slice(), so an int (a valid SliceInput) collapsed the neuron axis. Building the labels then raised "TypeError: iteration over a 0-d tensor", and that crash hid two more bugs: the non-LN path concatenated layers along positions (e.g. (8, 3, 16) instead of (2, 3, 4, 16)), and the LN-folded projection path raised IndexError (LN) or a broadcast RuntimeError (RMS). Normalise with Slice.unwrap() so an int keeps the neuron axis like [n] on every path. This makes the isinstance(neuron_labels, int) guard dead, so it is removed along with the now-unused numpy import. Add a regression test comparing neuron_slice=n against neuron_slice=[n] (labels, shape and values) for apply_ln on/off, no/vector/matrix project_output_onto, and incl_remainder on/off, on LN and RMS models. Fixes #1841 * Handle integer-mode Slice in stack_neuron_results Slice.unwrap() only converts a bare int into [n]; an existing Slice is returned unchanged. So neuron_slice=Slice(n) stayed in integer mode, collapsed the neuron axis, and building the labels raised "TypeError: iteration over a 0-d tensor", also with return_labels=False. Rebuild an integer-mode Slice as a new Slice([n]) inside stack_neuron_results only. The caller's Slice is not modified, and Slice / get_neuron_results keep their dimension-collapsing int semantics. Extend the regression test so Slice(n) is compared against [n] (labels, shape, values) over the same matrix as the bare int, and add a return_labels=False case for both forms.
* Fix per-layer head result cache validation * Validate cached head counts and cover batchless forward results
…ion (#1847) * Use the ln2 scale for neurons in fused LN+projection resid decomposition get_full_resid_decomposition(mlp_input=True, apply_ln=True, project_output_onto=p) normalised head, embed and bias rows with the ln2 scale but routed neurons through the fused LN+projection path, which always used the ln1 scale. The result therefore differed from the unprojected decomposition projected onto p and no longer summed to LN2(resid_mid) @ p. Expose mlp_input on stack_neuron_results (default False, so existing callers are unchanged), thread it into the fused helper and both apply_ln_to_stack calls, and pass it from get_full_resid_decomposition. Fixes #1846 * Fill the stack_neuron_results remainder up to resid_mid when mlp_input=True Follow the decompose_resid and accumulated_resid convention: with mlp_input=True the stack is the input to layer's MLP, so the remainder fills it up to resid_mid instead of resid_pre, and the stack sums to the (normalized) MLP input.
The expected score is baseline + shift, and the two partly cancel, so the float32 rounding error relative to the result can exceed pytest.approx's default rel=1e-6. With torch 2.14.1 on CPU the test fails at 1.3e-6 relative error (5.3e-7 with the locked torch 2.11). Use abs=1e-6, the tolerance the other edge-score comparisons in this file already use.
* Fix FactoredMatrix ellipsis indexing * Preserve nested Boolean sequence indexing in FactoredMatrix * Handle matrix new axes in FactoredMatrix indexing
* Fix per-example logit attribution across positions * fix: preserve shared per-position attribution targets --------- Co-authored-by: Nehal <nehalgajraj9@gmail.com> Co-authored-by: jlarson4 <jonahalarson@comcast.net>
…1854) * perf: reduce temporary memory when computing attention head results * perf: use contractions for attention head results * perf: use einsum throughout for cached attention head results
* fix: preserve attention masks in Inspect HF forwards * refactor(inspect): share position gate and declare mask capability
…#1859) * fix(bridge): slice pos before dropping batch dim in get_caching_hooks `BridgeCore.get_caching_hooks` removed the batch dim and then applied `pos_slice` on `_pos_slice_dim(name)`, which indexes the *batched* layout (dim 1 for resid / per-head tensors). With `remove_batch_dim=True` dim 1 is d_model or n_heads, so the cache held a slice of the wrong axis, e.g. `blocks.0.hook_out` -> [pos, 1] instead of [1, d_model]. A 2-D [batch, pos] activation (token ids at `embed.hook_in`) was not sliced at all, because the `dim() >= 2` guard ran on the 1-D remainder. Attention maps were unaffected because their axis is -2. Apply the slice first and drop the batch dim second, matching the order `run_with_cache._store` already uses, so both caching APIs agree. Adds a unit test (random-init `boot_native`, no downloads) that compares `get_caching_hooks` values against `run_with_cache` for int / tuple / list slices on token ids, a resid stream, a head-split tensor and an attention pattern. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(bridge): resolve the pos_slice axis per hook layout Reordering the slice before the batch drop regressed hooks whose position axis is not dim 1 in the batched layout: post-reshape qk-norm hooks (Gemma-3/Cohere, `[batch, heads, pos, d_head]`) and Mamba-1's channel-first conv1d / eager-scan hooks (`[batch, channels, pos, ...]`). `run_with_cache` sliced the same wrong axis for them. Add `HookPoint.pos_dim`, set by the component that fires on such a layout, and have `_pos_slice_dim` consult it for both caching APIs. Mark the post-reshape `hook_q_normed`/`hook_k_normed` and the `q_norm`/`k_norm` submodule hooks, `DepthwiseConv1DBridge.hook_in`/`hook_out`, and `SSMMixerBridge.hook_ssm_write`/`hook_ssm_state`. `hook_rot_q`/`hook_rot_k` already present `[batch, pos, heads, d_head]` via their conversion and are unchanged. Tests: value comparisons against the unsliced cache for a tiny Gemma-3 and a tiny Mamba-1 bridge, an `incl_bwd=True` gradient case, and trimmed comments. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(bridge): tag every head-major hook with its position axis Cover the remaining hooks whose batched layout does not carry position on dim 1, so pos_slice slices positions for both caching APIs: - GeneralizedComponent.mark_pos_dim() tags every HookPoint a component owns; post-reshape q_norm/k_norm now use it, which also covers the norm's own hook_normalized and hook_scale (Gemma-3, EXAONE-4, GLM-4-MoE, LFM2, Qwen3-Next). - MLAAttentionBridge fires hook_q/hook_k/hook_v after the head transpose ([batch, heads, pos, d]); tag them (DeepSeek-V3, GLM-MoE-DSA, GLM-4-MoE-Lite). hook_rot_q/hook_rot_k keep the default axis: TransposeRotaryHeads already hands the user [batch, pos, heads, d]. - T5Gemma2's hook_cross_pattern is [batch, heads, q_pos, k_pos] but its name misses the hook_pattern rule; tag it -2. Tests: the Gemma-3 case adds q_norm.hook_scale and k_norm.hook_normalized; a new DeepSeek-V3 / GLM-MoE-DSA case checks hook_q/k/v on dim 2 and the rotary hooks on dim 1 against manual indexing of the unsliced cache. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…motron adapters (#1869) * Set rmsnorm_uses_offset on the Qwen3.5 and Qwen3Next adapters HF's Qwen3.5, Qwen3.5-MoE and Qwen3Next RMSNorms scale by (1 + weight). Without the flag, NormalizationBridge dropped the 1 + whenever it computed a norm itself: on hook_scale/hook_normalized edits, with a backward hook attached, and in the LN-rule backward pass. Fixes #1868. Worked through with a hand from Claude. * Set rmsnorm_uses_offset on the Nemotron adapter NemotronLayerNorm1P applies gamma as (weight + 1). The adapter already disabled fold_ln for this, but never set the flag, so the bridge's own norm paths (hook_scale/hook_normalized edits, backward hooks, LN-rule) scaled by weight alone. * Keep rmsnorm_uses_offset off for Jais2 Jais2 subclasses the Nemotron adapter but uses plain nn.LayerNorm and re-enables fold_ln, so inheriting the offset made folding use 1 + w and broke compat-mode parity.
…mes (#1862) get_caching_hooks stored every activation under HookPoint.name, so a names_filter that matched an alias (e.g. blocks.0.hook_resid_pre) fired the hook but left no entry under the name the caller asked for, and cache[alias] raised KeyError. run_with_cache caches under both spellings. Collect every name the filter matched for a hook point and store the tensor (and its gradient) under the canonical name plus those aliases. The unfiltered sweep stays canonical-only. Co-authored-by: jlarson4 <jonahalarson@comcast.net>
… larger than 1 (#1865) * fix(bridge): keep the batch dimension in the caching hooks when it is larger than 1 remove_batch_dim=True on get_caching_hooks / add_caching_hooks took element 0 of every tensor, so a batch larger than 1 silently lost every example after the first. A hook sees one tensor at a time and cannot tell a batch dimension from a flattened [batch * pos, ...] one (MoE router and OPT ln2 hooks), so it now drops a leading dimension of size 1 and leaves every other tensor as it is. run_with_cache(return_cache_object=False) ignored the flag for a larger batch; it now makes the same batch-size check as the ActivationCache path. HookedRootModule gets the same hook rule. Its position-axis guard looked at the tensor before the batch dimension was dropped, so a [1, d] activation reached the slice as 1-D and raised IndexError; it now checks the stored tensor. utils.remove_batch_dim is annotated Float[Tensor, "1 ..."], which its runtime type check enforces, so it rejected every tensor it is documented to return unchanged. Annotate it as Tensor and let it accept a 0-d tensor. * fix(bridge): take the batch size from the input in run_with_cache(remove_batch_dim=True) ActivationCache.remove_batch_dim inferred the batch size from the most common leading dimension of the cached entries, so a filter that selects only flattened or position-indexed hooks (MoE router hooks, T5's pos_embed) made it read the sequence length as the batch size and raise at batch size 1. run_with_cache knows the real batch size from its input, so pass it in: remove_batch_dim takes an optional batch_size, used by both return types, and falls back to inferring it when the input does not state one. HookedRootModule.run_with_cache makes the same check on its first positional input, so both models raise for a batch larger than 1. Document that the caching hooks need no remove_batch_dim of their own under run_with_hooks(remove_batch_dim=True). --------- Co-authored-by: Jonah Larson <jonahalarson@comcast.net>
* fix: honor vLLM logits processor transforms in reconstruction * refactor(vllm): unify reconstruction cache and worker validation
…_dim=True) (#1864) * fix(bridge): let a hook return None under run_with_hooks(remove_batch_dim=True) The wrapper that hides the batch dimension from a hook called result.dim() on whatever the hook returned. A hook may return None to leave the activation alone, which is how a read-only hook is written, so at batch size 1 every such hook raised AttributeError. Only put the batch dimension back on a tensor result. * test(bridge): cover an in-place edit from a None-returning hook under remove_batch_dim A hook that returns None can still change the activation by editing it in place. That only reaches the model because squeeze(0) returns a view, so pin it with a case that fails if the wrapper ever hands the hook a copy. * test(bridge): pin that caching hooks under run_with_hooks(remove_batch_dim=True) need no flag Under the flag every hook already receives the tensor without its batch dimension, so the cache holds the same batch-free tensors as run_with_cache(remove_batch_dim=True). --------- Co-authored-by: jlarson4 <jonahalarson@comcast.net>
…#1873) CohereRotaryEmbedding returns cos/sin with each frequency repeated twice and Cohere rotates adjacent element pairs, but the attention bridge used the Llama formula, so logits did not match Hugging Face (hyper-accel/tiny-random-cohere: max diff 1.1e-2 and a different argmax before, 7.8e-8 after). Add a rotary_interleaved_cos_sin config flag, set by the Cohere adapter and inherited by Cohere2, and a helper that rotates adjacent pairs with cos/sin as given.
…1874) MPTALiBiAttentionBridge._reconstruct_attention never called _update_kv_cache, so each cached decoding step built K and V from the new token alone and the model ignored the prompt. Generation with the cache differed from Hugging Face (hf-internal-testing/tiny-random-MptForCausalLM: 8 of 26 ids), while generation without it matched. Call _update_kv_cache after the heads are split, as the Bloom and joint-QKV bridges do.
HunYuanDenseV1 rotates Q and K and then applies query_layernorm and key_layernorm, but the attention bridge normalised before RoPE for every post-reshape QK-norm, so logits did not match Hugging Face (tencent/Hunyuan-0.5B-Instruct: max diff 3.4e-2 before, 4.9e-5 after). Add a qk_norm_after_rope config flag, set by the HunYuan adapter, and a parity test against a tiny random HF model. Co-authored-by: Jonah Larson <jonahalarson@comcast.net>
* v4.2.0 PR follow ups, including cleanup of different model bridge components, and adding a few new tests for the new features. * notebook cleanup * test failure fix
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Release v4.2.0
Type of change
Checklist: