transformer_lens.model_bridge.supported_architectures.raven module

Raven / Huginn architecture adapter (RavenForCausalLM).

Family tomg-group-umd/huginn-0125: depth-recurrent (“latent reasoning”) decoder loaded via remote code. Three phases over one residual width — prelude blocks, a weight-tied recurrent core applied N times (num_steps is a RUNTIME forward argument, defaulting to config.mean_recurrence), then coda blocks. Core blocks are post-residual SandwichBlock modules (residual renormalised after each add), MHA with combined Wqkv plus a learned additive qk_bias.

Adapter decisions: - Full delegation: recurrence, prelude re-injection, emb-scale, sandwich

norms, and RoPE all run inside the remote-code forward.

  • OpaqueBlockBridge for all three block lists: BlockBridge’s hook aliases hardcode the standard pre-norm flow the SandwichBlock does not follow.

  • Core-block hook_in/hook_out fire once PER recurrence step (N times per forward); run_with_cache keeps the final step. Per-step access is not expressible through the static hook names — it needs the model’s native iterate_one_step/predict_from_latents interface. Deliberately not mapped: transformer.adapter and HuginnDynamicCache’s slot layout.

  • applicable_phases = []: initialize_state uses torch.randn_like, so the forward is non-deterministic unless seeded — integration tests pin the seed around both the bridge and HF calls.

  • Only bias anywhere is qk_bias; weight processing must tolerate missing biases via ProcessWeights._safe_get_tensor().

class transformer_lens.model_bridge.supported_architectures.raven.RavenArchitectureAdapter(cfg: Any)

Bases: ArchitectureAdapter

Architecture adapter for RavenForCausalLM (Huginn depth-recurrent decoder).

Prelude / weight-tied recurrent core / coda phases over a shared residual width. The recurrence and prelude re-injection live inside the remote-code HF forward, which the bridge delegates to; see the module docstring for the full set of adapter decisions.

__init__(cfg: Any) → None

Initialize the Raven / Huginn architecture adapter.

applicable_phases: list[int] = []
component_mapping: ComponentMapping | None
prepare_loading(model_name: str, model_kwargs: dict) → None

Patch Huginn’s remote code for transformers v5 compatibility.

Huginn’s modeling code targets transformers 4.44; two things break under v5 (5.8.1), so two patches:

  1. Tied-weights format. RavenForCausalLM._tied_weights_keys is a list (["lm_head.weight"], the 4.x format), but v5’s tie_weights -> get_expanded_tied_weights_keys calls .keys() on it and raises AttributeError. The model does not even construct. Rewrite it to the v5 dict form {"lm_head.weight": "transformer.wte.weight"} (Huginn ties lm_head to transformer.wte).

  2. Weight re-init. Under v5’s meta-device load-then-materialise flow, PreTrainedModel._init_weights is invoked on modules that already hold checkpoint weights, re-randomising them. Guard it to skip modules whose parameters are already on a real (non-meta) device — the same defensive patch openelm.py applies.

Parameters:
  • model_name – The HuggingFace model name/path.

  • model_kwargs – The kwargs dict for from_pretrained().

supports_batched_generation: bool = False
supports_kv_cache: bool = False
uses_split_attention: bool
weight_processing_conversions: Dict[str, ParamProcessingConversion | str] | None