transformer_lens.model_bridge.supported_architectures.qwen3_5 module¶
Qwen3.5 architecture adapter.
Hybrid linear-attention (GatedDeltaNet) + full-attention with dense gated MLP. 3 linear-attn layers per 1 full-attn layer. Extends Qwen3 base with optional attention mapping and fold_ln disabled.
- class transformer_lens.model_bridge.supported_architectures.qwen3_5.Qwen3_5ArchitectureAdapter(cfg: Any)¶
Bases:
Qwen3ArchitectureAdapterHybrid linear-attention + full-attention with dense gated MLP.
Inherits Qwen3 config/attention/MLP structure. Differences: - Attention + linear_attn are optional (per-layer type) - Gated q_proj: [query|gate] is split at forward time, never in weight space;
AttentionBridge exposes a query-only W_Q analysis view
- prepare_loading(model_name: str, model_kwargs: dict) None¶
Swap the multimodal config for its text_config so AutoModelForCausalLM loads the text-only model (published checkpoints carry the ForConditionalGeneration architecture).
- prepare_model(hf_model: Any) None¶
Reject full multimodal checkpoints on this text-only adapter.