Baimam Boukar

Understanding and Steering Language Models

I am interested in what we can learn from language models internals, and what that understanding lets us do. My current focus is the generalizability of interpretability methods across models, tasks and distribution shifts, and their implications for reliable interventions.

Baimam Boukar

Recent updates

Spoke at Ctrl+Alt+Échec
See all

Research In Progress

When does a probe generalize?

A reproduction-first audit across model families, with probes refit for each target model.