Skip to Main Content

Cross-Model Generalization of Mechanistic-Interpretability Probes

A reproduction-first audit of activation-probing methods across model families

Image will load when scrolled into view
Mechanistic InterpretabilityActivation ProbesGeneralizationReplication

About the Project

We study when activation-probing findings generalize across model families. The audit first reproduces selected methods under their original conditions, then evaluates the recovered findings on additional models.

Every target model receives a newly fitted probe using activations from that model; the language-model weights remain fixed. The experiment tests method generalization by independently refitting probes on each model.

Targeted experiments examine how response generation and labels, activation-site selection, probe architecture, and operating thresholds affect the scope of a generalization claim. I contributed experimental implementation, model runs, result analysis, and investigation of variation across evaluation conditions.

Project Details

StatusManuscript
Role
Research Assistant, Jinesis AI Lab
Stack
Python
PyTorch
Hugging Face Transformers
Activation Probing
Controlled Evaluation