Skip to content

fix(bridge): apply HunYuan's QK-norm after RoPE - #1872

Merged
jlarson4 merged 2 commits into
TransformerLensOrg:devfrom
Mudassiruddin7:fix-hunyuan-qk-norm-after-rope
Oct 9, 2026
Merged

jlarson4 merged 2 commits into
TransformerLensOrg:devfrom
Mudassiruddin7:fix-hunyuan-qk-norm-after-rope

Conversation

@Mudassiruddin7

Copy link
Copy Markdown
Contributor

Description

HunYuanDenseV1ForCausalLM does not match Hugging Face through the bridge, because the bridge normalises Q and K before RoPE and Hugging Face does it after. In modeling_hunyuan_v1_dense.py the attention forward calls apply_rotary_pos_emb and only then query_layernorm and key_layernorm. PositionEmbeddingsAttentionBridge applies a post-reshape QK-norm before RoPE, which is right for Gemma-3 and Cohere and wrong here.

This adds a config flag, qk_norm_after_rope (default False, declared like rotary_adjacent_pairs). When it is set, the bridge rotates first and then applies q_norm and k_norm; hook_rot_q and hook_rot_k still fire right after the rotation, and hook_q_normed and hook_k_normed fire after the norm. The HunYuan adapter sets it. Nothing else changes for the other architectures.

I found this with a parity battery over the supported causal-LM architectures (a tiny random-weight Hugging Face model per architecture wrapped with build_bridge_from_module, bridge against HF logits). In HunYuan the MLP outputs matched exactly and the attention output did not, which pointed at the attention order. The project's own record agrees: supported_models.json has tencent/Hunyuan-0.5B-Instruct below threshold with forward_pass_logits failing at max_diff=0.013494.

On the real checkpoint (fp32, CPU, a 13 token prompt, logit scale 16.4):

max abs diff to HF mean abs diff
dev 3.36e-02 2.80e-03
this PR 4.91e-05 4.49e-06

With a tiny random model with wide weights the difference goes from 5.2 to 5.5e-06. I also checked that the argmax agrees at every position, before and after. I did not run verify_models or the full benchmark.

Fixes # (no issue, found by the parity check above)

Type of change

  • Bug fix (non-breaking change which fixes an issue)

Tests

In tests/unit/model_bridge/supported_architectures/test_hunyuan_v1_dense_adapter.py: the adapter sets qk_norm_after_rope and the config default is False, and TestHunYuanDenseV1Parity builds a tiny random HunYuanDenseV1ForCausalLM (no download) and checks the bridge logits with assert_tiny_parity. All three fail on dev (the parity test with the drift assertion) and pass with this change. The file passes, 53 tests.

I ran all of tests/unit (Windows, Python 3.12, torch 2.14.1 CPU, transformers 5.19.0): 7220 passed. 46 tests failed or errored, the same 46 on unmodified dev. 43 of them need the optional timm and torchvision packages of the multimodal group, and pass once those are installed. The other 3 fail the same way with and without this change: test_model_structure_doc.py::test_every_qualified_hook_name_across_docs_exists (a UnicodeDecodeError reading a file with the Windows default codec) and two in test_sparse_probing.py (test_default_tolerance_accepts_large_scale_activations and test_stop_reason_distinguishes_converged_from_capped_fits), which I did not look into. I ran black, isort, pycln and mypy on the changed files. My black version also flags position_embeddings_attention.py on unmodified dev; I let it format that file, which only adds blank lines.

Checklist:

  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes (46 fail, identical on unmodified dev, see above)
  • I have not rewritten tests relating to key interfaces which would affect backward compatibility

HunYuanDenseV1 rotates Q and K and then applies query_layernorm and key_layernorm, but the attention bridge normalised before RoPE for every post-reshape QK-norm, so logits did not match Hugging Face (tencent/Hunyuan-0.5B-Instruct: max diff 3.4e-2 before, 4.9e-5 after). Add a qk_norm_after_rope config flag, set by the HunYuan adapter, and a parity test against a tiny random HF model.
Copilot AI balanced review requested due to automatic review settings October 9, 2026 13:48

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@jlarson4
jlarson4 merged commit 3f198fd into TransformerLensOrg:dev Oct 9, 2026
27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants