transformer_lens.model_bridge.supported_architectures.internlm2 module

InternLM2 architecture adapter.

class transformer_lens.model_bridge.supported_architectures.internlm2.InternLM2ArchitectureAdapter(cfg: Any)

Bases: ArchitectureAdapter

Architecture adapter for InternLM2 models.

InternLM2 uses remote code (trust_remote_code=True) and differs from Llama in: - Fused interleaved GQA wqkv weight (not standard [Q|K|V] split). The attention

bridge splits it at load and drops the fused key from the state dict, so fold_ln sees ordinary q/k/v keys and needs no special handling here.

  • Non-standard module names: tok_embeddings, output, attention, feed_forward, wqkv/wo, w1(gate)/w3(up)/w2(down), attention_norm, ffn_norm

  • Per-layer rotary_emb (no model-level shared instance)

Optional parameters (may not exist in state_dict): - blocks.{i}.attn.b_Q / b_K / b_V / b_O — config.bias=False on shipped models - blocks.{i}.mlp.b_gate / b_in / b_out — MLP always bias=False - blocks.{i}.ln1.b / ln2.b / ln_final.b — RMSNorm has no bias

prepare_loading(model_name: str, model_kwargs: dict) → None

Patch transformers v5 incompatibilities before from_pretrained runs.

prepare_model(hf_model: Any) → None

Restore per-layer rotary inv_freq lost to meta-device loading – this remote code predates HF’s original_inv_freq auto-restore, so positions would otherwise rotate by random values.

setup_component_testing(hf_model: Any, bridge_model: Any = None) → None

Inject per-layer rotary embedding for component testing.