Cross-Model Generalization of Mechanistic-Interpretability Probes
A reproduction-first audit of activation-probing methods across model families
About the Project
We study when activation-probing findings generalize across model families. The audit first reproduces selected methods under their original conditions, then evaluates the recovered findings on additional models.
Every target model receives a newly fitted probe using activations from that model; the language-model weights remain fixed. The experiment tests method generalization by independently refitting probes on each model.
Targeted experiments examine how response generation and labels, activation-site selection, probe architecture, and operating thresholds affect the scope of a generalization claim. I contributed experimental implementation, model runs, result analysis, and investigation of variation across evaluation conditions.