transformer_lens.model_bridge.supported_architectures.apertus module

Apertus architecture adapter.

class transformer_lens.model_bridge.supported_architectures.apertus.ApertusArchitectureAdapter(cfg: Any)

Bases: ArchitectureAdapter

Architecture adapter for Apertus models.

Apertus uses a pre-norm architecture with RMSNorm, Q/K normalization in attention, rotary position embeddings (RoPE with LLaMA-3 scaling), grouped query attention (GQA), non-gated MLP (XiELU activation), and no biases on any projections.

Similar to Qwen3 (pre-norm RMSNorm, QK-norm, GQA, RoPE) but uses a non-gated MLP (up_proj -> XiELU -> down_proj) instead of gated MLP.

Note: Apertus uses different layer norm names than most Llama-family models: - attention_layernorm (instead of input_layernorm) - feedforward_layernorm (instead of post_attention_layernorm)

__init__(cfg: Any) None

Initialize the Apertus architecture adapter.

prepare_loading(model_name: str, model_kwargs: dict) None

Patch XIELUActivation to defer eager .item() calls for meta tensor compat.

Transformers v5 uses meta tensors during from_pretrained, but XIELUActivation.__init__ eagerly calls .item() on beta/eps buffers to precompute _beta_scalar/_eps_scalar for the CUDA kernel path. This fails on meta device. Once upstream fixes this (transformers PR #43473), this patch can be removed.

Instead of reimplementing __init__, we wrap it to catch the meta tensor failure and defer scalar computation to forward() time.