transformer_lens.tools.model_registry.registry_io module

Shared I/O functions for reading and writing model registry data files.

Consolidates the load-modify-save pattern used by verify_models.py and main_benchmark.py into a single module that properly uses the VerificationRecord/VerificationHistory dataclasses.

transformer_lens.tools.model_registry.registry_io.add_verification_record(model_id: str, arch_id: str, notes: str | None = None, verified_by: str = 'verify_models', sanitize_fn: Callable[[str | None], str | None] | None = None, prompt_profile: str | None = None, p4_scoring_version: int | None = None) → None

Append a VerificationRecord to verification_history.json.

Uses the VerificationRecord dataclass properly instead of raw dict manipulation.

Parameters:
  • model_id – The verified model

  • arch_id – Architecture type

  • notes – Optional verification notes

  • verified_by – Who/what performed the verification

  • sanitize_fn – Optional callable to sanitize note strings

transformer_lens.tools.model_registry.registry_io.extract_phase_scores(results: list) → dict[int, float | None]

Extract phase scores from benchmark results.

Shared home for both registry-writing paths (verify_models and main_benchmark.update_model_registry) so they cannot drift.

Parameters:

results – List of BenchmarkResult objects

Returns:

Dict mapping phase number to score (0-100) or None

transformer_lens.tools.model_registry.registry_io.is_hf_loadable_quantized(model_id: str) → bool

True for quantizations loadable by HF transformers + a quant library.

transformer_lens.tools.model_registry.registry_io.is_incompatible_quantized(model_id: str) → bool

True for quantization formats the bridge can’t ingest (GGUF, MLX, FP4/FP8).

transformer_lens.tools.model_registry.registry_io.is_quantized_model(model_id: str) → bool

Alias for is_incompatible_quantized — kept for back-compat with existing call sites.

transformer_lens.tools.model_registry.registry_io.load_model_aliases() → dict[str, list[str]]

Load the canonical alias table: official HF model name -> deprecated short aliases.

transformer_lens.tools.model_registry.registry_io.load_supported_models_raw() → dict

Load supported_models.json as a raw dict.

transformer_lens.tools.model_registry.registry_io.load_verification_history() → VerificationHistory

Load verification_history.json into a VerificationHistory dataclass.

transformer_lens.tools.model_registry.registry_io.pass_status(use_hf_reference: bool) → int

Status for a passing run: VERIFIED with an HF reference, else PROVISIONAL (a –no-hf-reference structural-only pass is recorded but not counted verified).

transformer_lens.tools.model_registry.registry_io.recompute_registry_totals(models: list[dict]) → dict

Header totals for supported_models.json, recomputed from the models list.

Shared by both writers (update_model_status here and hf_scraper’s report builder) so the counting rules cannot drift.

transformer_lens.tools.model_registry.registry_io.registry_prompt_profile(model_id: str) → str | None

Stored prompt_profile for a model, or None. Uncached read: the sweep rewrites the registry between models.

transformer_lens.tools.model_registry.registry_io.required_quant_library_for_model(model_id: str) → str | None

Return the Python import name needed to load this model, or None if unquantized.

transformer_lens.tools.model_registry.registry_io.resolve_model_alias(model_name: str) → str | None

Return the official HF name if model_name is a deprecated alias, else None.

transformer_lens.tools.model_registry.registry_io.save_supported_models_raw(data: dict) → None

Save raw dict back to supported_models.json.

transformer_lens.tools.model_registry.registry_io.save_verification_history(history: VerificationHistory) → None

Save VerificationHistory dataclass to verification_history.json.

transformer_lens.tools.model_registry.registry_io.update_model_status(model_id: str, arch_id: str, status: int | None = None, note: str | None = None, phase_scores: dict[int, float | None] | None = None, sanitize_fn: Callable[[str | None], str | None] | None = None, prompt_profile: str | None = None) → bool

Update a single model entry in supported_models.json.

If the model is not found in the registry and status is STATUS_VERIFIED or STATUS_PROVISIONAL, a new entry is appended.

When status is None (partial-phase update), only the provided phase_scores are updated — status, note, and other scores are preserved.

Parameters:
  • model_id – The model to update

  • arch_id – Architecture of the model

  • status – New status code (0-4), or None for score-only updates

  • note – Optional note for skip/fail reason

  • phase_scores – Phase score dict {1: float, 2: float, 3: float, 4: float}

  • sanitize_fn – Optional callable to sanitize note strings

  • prompt_profile – Phase-4 prompt profile actually used (e.g. “task:translation@en-de”). Sparse: the default “continuation” removes the key (clearing a stale non-default value), None (no Phase-4 result) leaves it untouched.

Returns:

True if entry was found/created and updated