Skip to content

Release v4.2.0 - #1876

Merged
jlarson4 merged 36 commits into
mainfrom
dev
Oct 10, 2026
Merged

jlarson4 merged 36 commits into
mainfrom
dev

Conversation

@jlarson4

@jlarson4 jlarson4 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

Description

Release v4.2.0

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • This change requires a documentation update

Checklist:

  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • I have not rewritten tests relating to key interfaces which would affect backward compatibility

MdSadiqMd and others added 30 commits September 28, 2026 09:48
…l-seeding

fix sparse probe control seeding
…1820)

* Add tie-aware roc_auc and average_precision to sparse probe metrics

F1 and the other threshold metrics are read at logit >= 0 only, so a probe
on a dead (constant) coordinate that predicts every held-out example positive
scores F1 = 2p/(1+p). Add threshold-free, tie-aware ROC-AUC (Mann-Whitney
with average ranks) and average precision (tied logits form one threshold)
to SparseProbeMetrics, propagate both to SparseProbeControl, and document
the degenerate case in the sparse probing guide.

Fixes #1814

* Clarify the single-class policy for threshold-free probe metrics
…mpute dtype (#1829)

* test(projection_kernel): reproduce the half-precision rank collapse

The default relative rank tolerance is floored at the storage dtype's epsilon.
For bfloat16 storage that floor is about 7.8e-3, which is far above the
singular-value ratios that separate a full-rank head from a numerically
rank-deficient one. A full-rank head whose spectrum decays gradually therefore
collapses to a much smaller measured rank, and a caller requesting the full
width fails with "requested rank N exceeds measured rank M".

Add a fixture that builds a full-rank matrix with a geometric singular-value
spectrum, plus a single-head bridge that injects it as W_Q. Cover the collapse
at both half-precision storage dtypes, confirm a structurally rank-deficient
half-precision matrix is still detected, and pin the float32 path as unchanged
so the tests distinguish half-precision storage from genuinely deficient data.

These tests fail against the current default and pass once the tolerance is
derived from the compute dtype.

* fix(projection_kernel): derive the default rank tolerance from the compute dtype

The default relative rank tolerance was floored at the storage dtype's epsilon.
For bfloat16 storage that floor is about 7.8e-3, which is far above the
singular-value ratios that separate a full-rank head from a numerically
rank-deficient one. A full-rank head whose spectrum decays gradually therefore
collapsed to a much smaller measured rank, and a caller requesting the full
width failed with "requested rank N exceeds measured rank M".

The SVD already promotes float16 and bfloat16 inputs to float32, so the measured
singular values carry float32 precision regardless of how the checkpoint was
stored. Flooring the tolerance at the storage epsilon discarded that precision
and made the rank decision depend on the storage dtype rather than on the
numerics actually computed.

Derive the default from the compute dtype instead, and drop the storage-dtype
helper that existed only to feed the old floor. An explicit rtol still takes
precedence.

One existing deficiency fixture relied on the loose floor to register a
near-degenerate direction; it now uses an exactly dependent column so it tests
structural deficiency rather than quantization noise.

* docs(projection_kernel): describe the compute-dtype default rank tolerance

The guide still documented the default relative rank tolerance as
`max(max(matrix.shape) * compute_eps, storage_eps)`, which no longer matches the
implementation. Rewrite both passages to describe the compute-dtype default and
say why it is derived from the dtype the SVD runs in rather than the storage
dtype: half-precision inputs are promoted to float32, so a storage-dtype floor
would discard the precision of the singular values actually computed and
collapse the measured rank of a full-rank half-precision head.

Also note that an explicit rtol still takes precedence, and that half-precision
checkpoints no longer need a manual override.

* test(projection_kernel): cover the half-precision wrapper path on gpt2-small

Add an integration case that loads gpt2-small with bfloat16 weights and runs the
default O -> Q call without an explicit rtol. gpt2-small is dense MHA with no GQA
grouping, so this isolates the storage dtype effect from the head-cardinality
effect.

The discriminating assertion is the effective tolerance: it must sit below the
bfloat16 epsilon, which is what the storage-dtype floor violated. The measured
ranks are additionally pinned against the float32 run, and the scores are checked
for finiteness and the documented bounds. Exact score equality across dtypes is
deliberately not asserted.

gpt2-small's worst-conditioned head sits near 40, below the reciprocal of the
bfloat16 epsilon, so its measured ranks survive even the storage-dtype floor.
The rank and bound checks therefore guard the end-to-end contract rather than
reproducing the collapse; a head whose spectrum decays past that reciprocal is
covered by the synthetic unit cases.
Co-authored-by: jlarson4 <jonahalarson@comcast.net>
* Forward the requested revision to every remote-code force-import

prepare_loading force-imports remote modeling classes so their modules
land in sys.modules to patch. Each Hub revision is its own module copy,
and only the default revision was imported, so a pinned
revision=... loaded by from_pretrained afterwards got an unpatched copy.
#1810 fixed this for restore_default_rope_init; do the same for Dream's
DreamGenerationConfig and DreamAttention imports and for the BD3LM,
GIDD, InternLM2, OpenELM, Ouro, Raven and RWKV-7 adapters.

* Forward the requested revision in Baichuan's force-import
* fix(sparse_probing): add std_floor to standardization

Near-constant columns were being divided by their own tiny std,
blowing noise up to unit scale. The scale now has a floor
(default 1e-3, same as the reference). Zero-variance columns
still get scale 1, and std_floor=0 keeps the old behavior.

Also document the max_eval budget in the max_iter docstrings.

* sparse_probing: explain zero-variance scale rule and test setup

* sparse_probing: address review feedback

Test that sweep_sparse_probe passes std_floor to its main and
control fits, fix the zero-variance guard comment, and note that
the line search can go one evaluation past the max_eval budget.

* sparse_probing: check sweep's default std_floor
* Add Muse Glimmer architecture adapter

Bridge MuseGlimmerForConditionalGeneration (text path through model.language_model and lm_head, vision tower and projector as opaque components), register it in the factory, model registry, multimodal list and HF model_type map, and pass output_multiplier through to the bridge config. The unit test compares bridge logits and attention patterns against HF on a tiny random config and skips when transformers lacks muse_glimmer.

* Pin the Muse Glimmer output-logit contract

The tiny Muse Glimmer model's logits are too small for the tanh cap to change
anything, so the forward-parity test cannot catch a missing post-unembedding
transform. Add a contract case alongside Granite and Falcon-H1:

- nested text_config: final_logit_softcapping and output_multiplier both reach
  the bridge config;
- apply_output_logits_transform applies the multiplier before the softcap
  (20 * tanh(0.5 * logits / 20)), with logits large enough that dropping the
  factor fails the comparison.

* Expose the merged vision embedding on Muse Glimmer's vision projector

HF runs vision_adapter -> vision_projection -> perception_emb_norm and scatters
only the last value into the text embeddings, so the vision_projector bridge now
wraps perception_emb_norm instead of the raw vision_projection Linear. Its
hook_out is then the embedding HF actually merges, matching the Gemma 3 and LLaVA
adapters; hook_in is the projection output.

Add a structural assertion to the adapter test.
pyproject allows beartype>=0.14.1, but uv.lock still pins 0.14.1. Newer
beartype (0.22.x, what a fresh install resolves to) checks dict values and
container items, and the jaxtyping test hook then rejects three call sites:

- TLWorkerExtension.tl_read_batched_captures returns encode_tensor() dicts,
  not tensors; annotate it like tl_read_captures.
- select_displacement_matched_control_token is called with the
  decomposition's support tensor; accept a tensor for active_support.
- test_resolve_state_dict_key_dense_mlp_fallback passed ints as state-dict
  values; use tensors.
* Fix recursive torch dispatch for CompositionScores sequences

* Consolidate CompositionScores regression tests in protocol class
* wire LLaVA vision hook aliases

* fix(vision): preserve CLIP layer norm epsilon
* fix(bridge): preserve seq2seq state in streaming generation

* fix(bridge): align seq2seq stream setup and stopping guards
…1845)

DynamicCache(config=...) on newer transformers (5.17) compares
config.num_kv_shared_layers against an int, which fails on a bare
MagicMock config. Use a tiny real NemotronHConfig instead.
* Fix stack_neuron_results for an integer neuron_slice

stack_neuron_results wrapped neuron_slice with Slice(), so an int (a valid
SliceInput) collapsed the neuron axis. Building the labels then raised
"TypeError: iteration over a 0-d tensor", and that crash hid two more bugs:
the non-LN path concatenated layers along positions (e.g. (8, 3, 16) instead
of (2, 3, 4, 16)), and the LN-folded projection path raised IndexError (LN)
or a broadcast RuntimeError (RMS).

Normalise with Slice.unwrap() so an int keeps the neuron axis like [n] on
every path. This makes the isinstance(neuron_labels, int) guard dead, so it
is removed along with the now-unused numpy import.

Add a regression test comparing neuron_slice=n against neuron_slice=[n]
(labels, shape and values) for apply_ln on/off, no/vector/matrix
project_output_onto, and incl_remainder on/off, on LN and RMS models.

Fixes #1841

* Handle integer-mode Slice in stack_neuron_results

Slice.unwrap() only converts a bare int into [n]; an existing Slice is
returned unchanged. So neuron_slice=Slice(n) stayed in integer mode,
collapsed the neuron axis, and building the labels raised
"TypeError: iteration over a 0-d tensor", also with return_labels=False.

Rebuild an integer-mode Slice as a new Slice([n]) inside
stack_neuron_results only. The caller's Slice is not modified, and Slice /
get_neuron_results keep their dimension-collapsing int semantics.

Extend the regression test so Slice(n) is compared against [n] (labels,
shape, values) over the same matrix as the bare int, and add a
return_labels=False case for both forms.
* Fix per-layer head result cache validation

* Validate cached head counts and cover batchless forward results
…ion (#1847)

* Use the ln2 scale for neurons in fused LN+projection resid decomposition

get_full_resid_decomposition(mlp_input=True, apply_ln=True,
project_output_onto=p) normalised head, embed and bias rows with the ln2
scale but routed neurons through the fused LN+projection path, which
always used the ln1 scale. The result therefore differed from the
unprojected decomposition projected onto p and no longer summed to
LN2(resid_mid) @ p.

Expose mlp_input on stack_neuron_results (default False, so existing
callers are unchanged), thread it into the fused helper and both
apply_ln_to_stack calls, and pass it from get_full_resid_decomposition.

Fixes #1846

* Fill the stack_neuron_results remainder up to resid_mid when mlp_input=True

Follow the decompose_resid and accumulated_resid convention: with mlp_input=True the stack is the input to layer's MLP, so the remainder fills it up to resid_mid instead of resid_pre, and the stack sums to the (normalized) MLP input.
The expected score is baseline + shift, and the two partly cancel, so the
float32 rounding error relative to the result can exceed pytest.approx's
default rel=1e-6. With torch 2.14.1 on CPU the test fails at 1.3e-6
relative error (5.3e-7 with the locked torch 2.11). Use abs=1e-6, the
tolerance the other edge-score comparisons in this file already use.
* Fix FactoredMatrix ellipsis indexing

* Preserve nested Boolean sequence indexing in FactoredMatrix

* Handle matrix new axes in FactoredMatrix indexing
* Fix per-example logit attribution across positions

* fix: preserve shared per-position attribution targets

---------

Co-authored-by: Nehal <nehalgajraj9@gmail.com>
Co-authored-by: jlarson4 <jonahalarson@comcast.net>
…1854)

* perf: reduce temporary memory when computing attention head results

* perf: use contractions for attention head results

* perf: use einsum throughout for cached attention head results
* fix: preserve attention masks in Inspect HF forwards

* refactor(inspect): share position gate and declare mask capability
…#1859)

* fix(bridge): slice pos before dropping batch dim in get_caching_hooks

`BridgeCore.get_caching_hooks` removed the batch dim and then applied
`pos_slice` on `_pos_slice_dim(name)`, which indexes the *batched* layout
(dim 1 for resid / per-head tensors). With `remove_batch_dim=True` dim 1 is
d_model or n_heads, so the cache held a slice of the wrong axis, e.g.
`blocks.0.hook_out` -> [pos, 1] instead of [1, d_model]. A 2-D [batch, pos]
activation (token ids at `embed.hook_in`) was not sliced at all, because the
`dim() >= 2` guard ran on the 1-D remainder. Attention maps were unaffected
because their axis is -2.

Apply the slice first and drop the batch dim second, matching the order
`run_with_cache._store` already uses, so both caching APIs agree.

Adds a unit test (random-init `boot_native`, no downloads) that compares
`get_caching_hooks` values against `run_with_cache` for int / tuple / list
slices on token ids, a resid stream, a head-split tensor and an attention
pattern.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(bridge): resolve the pos_slice axis per hook layout

Reordering the slice before the batch drop regressed hooks whose position
axis is not dim 1 in the batched layout: post-reshape qk-norm hooks
(Gemma-3/Cohere, `[batch, heads, pos, d_head]`) and Mamba-1's channel-first
conv1d / eager-scan hooks (`[batch, channels, pos, ...]`). `run_with_cache`
sliced the same wrong axis for them.

Add `HookPoint.pos_dim`, set by the component that fires on such a layout,
and have `_pos_slice_dim` consult it for both caching APIs. Mark the
post-reshape `hook_q_normed`/`hook_k_normed` and the `q_norm`/`k_norm`
submodule hooks, `DepthwiseConv1DBridge.hook_in`/`hook_out`, and
`SSMMixerBridge.hook_ssm_write`/`hook_ssm_state`. `hook_rot_q`/`hook_rot_k`
already present `[batch, pos, heads, d_head]` via their conversion and are
unchanged.

Tests: value comparisons against the unsliced cache for a tiny Gemma-3 and a
tiny Mamba-1 bridge, an `incl_bwd=True` gradient case, and trimmed comments.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(bridge): tag every head-major hook with its position axis

Cover the remaining hooks whose batched layout does not carry position on
dim 1, so pos_slice slices positions for both caching APIs:

- GeneralizedComponent.mark_pos_dim() tags every HookPoint a component owns;
  post-reshape q_norm/k_norm now use it, which also covers the norm's own
  hook_normalized and hook_scale (Gemma-3, EXAONE-4, GLM-4-MoE, LFM2, Qwen3-Next).
- MLAAttentionBridge fires hook_q/hook_k/hook_v after the head transpose
  ([batch, heads, pos, d]); tag them (DeepSeek-V3, GLM-MoE-DSA, GLM-4-MoE-Lite).
  hook_rot_q/hook_rot_k keep the default axis: TransposeRotaryHeads already
  hands the user [batch, pos, heads, d].
- T5Gemma2's hook_cross_pattern is [batch, heads, q_pos, k_pos] but its name
  misses the hook_pattern rule; tag it -2.

Tests: the Gemma-3 case adds q_norm.hook_scale and k_norm.hook_normalized;
a new DeepSeek-V3 / GLM-MoE-DSA case checks hook_q/k/v on dim 2 and the
rotary hooks on dim 1 against manual indexing of the unsliced cache.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…motron adapters (#1869)

* Set rmsnorm_uses_offset on the Qwen3.5 and Qwen3Next adapters

HF's Qwen3.5, Qwen3.5-MoE and Qwen3Next RMSNorms scale by (1 + weight).
Without the flag, NormalizationBridge dropped the 1 + whenever it computed
a norm itself: on hook_scale/hook_normalized edits, with a backward hook
attached, and in the LN-rule backward pass.

Fixes #1868.

Worked through with a hand from Claude.

* Set rmsnorm_uses_offset on the Nemotron adapter

NemotronLayerNorm1P applies gamma as (weight + 1). The adapter already
disabled fold_ln for this, but never set the flag, so the bridge's own
norm paths (hook_scale/hook_normalized edits, backward hooks, LN-rule)
scaled by weight alone.

* Keep rmsnorm_uses_offset off for Jais2

Jais2 subclasses the Nemotron adapter but uses plain nn.LayerNorm and
re-enables fold_ln, so inheriting the offset made folding use 1 + w and
broke compat-mode parity.
…mes (#1862)

get_caching_hooks stored every activation under HookPoint.name, so a
names_filter that matched an alias (e.g. blocks.0.hook_resid_pre) fired
the hook but left no entry under the name the caller asked for, and
cache[alias] raised KeyError. run_with_cache caches under both spellings.

Collect every name the filter matched for a hook point and store the
tensor (and its gradient) under the canonical name plus those aliases.
The unfiltered sweep stays canonical-only.

Co-authored-by: jlarson4 <jonahalarson@comcast.net>
… larger than 1 (#1865)

* fix(bridge): keep the batch dimension in the caching hooks when it is larger than 1

remove_batch_dim=True on get_caching_hooks / add_caching_hooks took element 0 of
every tensor, so a batch larger than 1 silently lost every example after the
first. A hook sees one tensor at a time and cannot tell a batch dimension from
a flattened [batch * pos, ...] one (MoE router and OPT ln2 hooks), so it now drops
a leading dimension of size 1 and leaves every other tensor as it is.

run_with_cache(return_cache_object=False) ignored the flag for a larger batch;
it now makes the same batch-size check as the ActivationCache path.

HookedRootModule gets the same hook rule. Its position-axis guard looked at the
tensor before the batch dimension was dropped, so a [1, d] activation reached
the slice as 1-D and raised IndexError; it now checks the stored tensor.

utils.remove_batch_dim is annotated Float[Tensor, "1 ..."], which its runtime
type check enforces, so it rejected every tensor it is documented to return
unchanged. Annotate it as Tensor and let it accept a 0-d tensor.

* fix(bridge): take the batch size from the input in run_with_cache(remove_batch_dim=True)

ActivationCache.remove_batch_dim inferred the batch size from the most common
leading dimension of the cached entries, so a filter that selects only flattened
or position-indexed hooks (MoE router hooks, T5's pos_embed) made it read the
sequence length as the batch size and raise at batch size 1.

run_with_cache knows the real batch size from its input, so pass it in: remove_batch_dim
takes an optional batch_size, used by both return types, and falls back to inferring it
when the input does not state one.

HookedRootModule.run_with_cache makes the same check on its first positional input, so both
models raise for a batch larger than 1. Document that the caching hooks need no
remove_batch_dim of their own under run_with_hooks(remove_batch_dim=True).

---------

Co-authored-by: Jonah Larson <jonahalarson@comcast.net>
* fix: honor vLLM logits processor transforms in reconstruction

* refactor(vllm): unify reconstruction cache and worker validation
Mudassiruddin7 and others added 6 commits October 8, 2026 16:48
…_dim=True) (#1864)

* fix(bridge): let a hook return None under run_with_hooks(remove_batch_dim=True)

The wrapper that hides the batch dimension from a hook called result.dim() on
whatever the hook returned. A hook may return None to leave the activation
alone, which is how a read-only hook is written, so at batch size 1 every such
hook raised AttributeError.

Only put the batch dimension back on a tensor result.

* test(bridge): cover an in-place edit from a None-returning hook under remove_batch_dim

A hook that returns None can still change the activation by editing it in
place. That only reaches the model because squeeze(0) returns a view, so pin it
with a case that fails if the wrapper ever hands the hook a copy.

* test(bridge): pin that caching hooks under run_with_hooks(remove_batch_dim=True) need no flag

Under the flag every hook already receives the tensor without its batch dimension,
so the cache holds the same batch-free tensors as run_with_cache(remove_batch_dim=True).

---------

Co-authored-by: jlarson4 <jonahalarson@comcast.net>
…#1873)

CohereRotaryEmbedding returns cos/sin with each frequency repeated twice and Cohere rotates adjacent element pairs, but the attention bridge used the Llama formula, so logits did not match Hugging Face (hyper-accel/tiny-random-cohere: max diff 1.1e-2 and a different argmax before, 7.8e-8 after). Add a rotary_interleaved_cos_sin config flag, set by the Cohere adapter and inherited by Cohere2, and a helper that rotates adjacent pairs with cos/sin as given.
…1874)

MPTALiBiAttentionBridge._reconstruct_attention never called _update_kv_cache, so each cached decoding step built K and V from the new token alone and the model ignored the prompt. Generation with the cache differed from Hugging Face (hf-internal-testing/tiny-random-MptForCausalLM: 8 of 26 ids), while generation without it matched. Call _update_kv_cache after the heads are split, as the Bloom and joint-QKV bridges do.
HunYuanDenseV1 rotates Q and K and then applies query_layernorm and key_layernorm, but the attention bridge normalised before RoPE for every post-reshape QK-norm, so logits did not match Hugging Face (tencent/Hunyuan-0.5B-Instruct: max diff 3.4e-2 before, 4.9e-5 after). Add a qk_norm_after_rope config flag, set by the HunYuan adapter, and a parity test against a tiny random HF model.

Co-authored-by: Jonah Larson <jonahalarson@comcast.net>
* v4.2.0 PR follow ups, including cleanup of different model bridge components, and adding a few new tests for the new features.

* notebook cleanup

* test failure fix
@jlarson4
jlarson4 merged commit bd55eab into main Oct 10, 2026
104 of 108 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.