transformer_lens.model_bridge.supported_architectures.raven module¶
Raven / Huginn architecture adapter (RavenForCausalLM).
Family tomg-group-umd/huginn-0125: depth-recurrent (“latent reasoning”)
decoder loaded via remote code. Three phases over one residual width —
prelude blocks, a weight-tied recurrent core applied N times (num_steps
is a RUNTIME forward argument, defaulting to config.mean_recurrence),
then coda blocks. Core blocks are post-residual SandwichBlock modules
(residual renormalised after each add), MHA with combined Wqkv plus a
learned additive qk_bias.
Adapter decisions: - Full delegation: recurrence, prelude re-injection, emb-scale, sandwich
norms, and RoPE all run inside the remote-code forward.
OpaqueBlockBridgefor all three block lists: BlockBridge’s hook aliases hardcode the standard pre-norm flow the SandwichBlock does not follow.Core-block
hook_in/hook_outfire once PER recurrence step (N times per forward);run_with_cachekeeps the final step. Per-step access is not expressible through the static hook names — it needs the model’s nativeiterate_one_step/predict_from_latentsinterface. Deliberately not mapped:transformer.adapterandHuginnDynamicCache’s slot layout.applicable_phases = []:initialize_stateusestorch.randn_like, so the forward is non-deterministic unless seeded — integration tests pin the seed around both the bridge and HF calls.Only bias anywhere is
qk_bias; weight processing must tolerate missing biases viaProcessWeights._safe_get_tensor().
- class transformer_lens.model_bridge.supported_architectures.raven.RavenArchitectureAdapter(cfg: Any)¶
Bases:
ArchitectureAdapterArchitecture adapter for RavenForCausalLM (Huginn depth-recurrent decoder).
Prelude / weight-tied recurrent core / coda phases over a shared residual width. The recurrence and prelude re-injection live inside the remote-code HF forward, which the bridge delegates to; see the module docstring for the full set of adapter decisions.
- __init__(cfg: Any) None¶
Initialize the Raven / Huginn architecture adapter.
- applicable_phases: list[int] = []¶
- component_mapping: ComponentMapping | None¶
- prepare_loading(model_name: str, model_kwargs: dict) None¶
Patch Huginn’s remote code for transformers v5 compatibility.
Huginn’s modeling code targets transformers 4.44; two things break under v5 (5.8.1), so two patches:
Tied-weights format.
RavenForCausalLM._tied_weights_keysis a list (["lm_head.weight"], the 4.x format), but v5’stie_weights->get_expanded_tied_weights_keyscalls.keys()on it and raisesAttributeError. The model does not even construct. Rewrite it to the v5 dict form{"lm_head.weight": "transformer.wte.weight"}(Huginn tieslm_headtotransformer.wte).Weight re-init. Under v5’s meta-device load-then-materialise flow,
PreTrainedModel._init_weightsis invoked on modules that already hold checkpoint weights, re-randomising them. Guard it to skip modules whose parameters are already on a real (non-meta) device — the same defensive patch openelm.py applies.
- Parameters:
model_name – The HuggingFace model name/path.
model_kwargs – The kwargs dict for from_pretrained().
- supports_batched_generation: bool = False¶
- supports_kv_cache: bool = False¶
- uses_split_attention: bool¶
- weight_processing_conversions: Dict[str, ParamProcessingConversion | str] | None¶