Work

Research, experiments, and the software behind them.

Cross-Model Generalization of Mechanistic-Interpretability Probes

Manuscript

A reproduction-first audit across model families, with probes refit for each target model.

Mechanistic Interpretability · Activation Probes

Selection Regret: Auditing Deception Probes as Intervention Selectors

Under Review

Testing whether a probe that detects deception can also select an intervention that changes behavior.

Mechanistic Interpretability · Deception Monitoring

Coordinate or Concept Mismatch?

Ongoing Research

Investigating cross-model probe-transfer failures and task-conditioned alignment

Mechanistic Interpretability · Probe Transfer

Measuring Delusion Reinforcement Across the Directness-Need Spectrum

Ongoing Research

Separating affective support from epistemic sycophancy in conversational models

LLM Safety · Epistemic Sycophancy

Replicating Matryoshka Sparse Autoencoders

Published Writeup (LessWrong)

Replicating feature recovery and comparing Matryoshka and standard sparse autoencoders

Mechanistic Interpretability · Sparse Autoencoders

Archive