Skip to Main Content

Hi, I'm Baimam Boukar 👋

I study when interpretability methods generalize and when internal model signals can guide reliable interventions. My work focuses on activation probes, steering, and evaluation across models, tasks, and distribution shifts.

Profile

Recent updates

Sep 20, 2026 Spoke at Ctrl+Alt+Échec

See All

Research Focus

Generalization, reliable monitoring, and activation-level interventions.

Mechanistic Interpretability

When probing findings remain reliable across changes in models, layers, tasks, and data.

Reliable Monitoring

Separating predictive accuracy from evidence about effective behavioral intervention.

Activation-Level Interventions

Testing steering directions and layers while preserving unrelated capabilities.

Controlled Evaluation

Reproduction, held-out tests, and explicit operating thresholds under distribution shift.


Selected Work

View All

Built from scratch with ••