transformer_lens.model_bridge.supported_architectures.qwen3_5_moe module

Qwen3.5-MoE architecture adapter.

Hybrid linear-attention (GatedDeltaNet) + full-attention with sparse MoE MLP (256 experts, top-8 routing, shared expert in public checkpoints). Same hybrid design as Qwen3.5 dense and the same MoE block family as Qwen3-Next.

Two adapters: text-only Qwen3_5MoeForCausalLM and the vision-language Qwen3_5MoeForConditionalGeneration (text backbone nested under model.language_model plus the Qwen3.5 vision tower).

class transformer_lens.model_bridge.supported_architectures.qwen3_5_moe.Qwen3_5MoeArchitectureAdapter(cfg: Any)

Bases: Qwen3_5ArchitectureAdapter

Text-only Qwen3.5-MoE: hybrid GatedDeltaNet + full attention, sparse MoE MLP.

class transformer_lens.model_bridge.supported_architectures.qwen3_5_moe.Qwen3_5MoeMultimodalArchitectureAdapter(cfg: Any)

Bases: Qwen3_5MultimodalArchitectureAdapter

Vision-language adapter for Qwen3_5MoeForConditionalGeneration.

Reuses the Qwen3.5 multimodal wiring (language model under model.language_model + vision tower) with the MLP swapped for sparse MoE.