How I am building my foundations in AI safety research — through mentored research, structured training, and hands-on replications.
Research Assistant — Jinesis AI Lab, University of Toronto
June 2026 - Present
Researching the generalizability of mechanistic interpretability techniques: when do results established on one model transfer to other models, tasks, and settings?
BlueDot Impact — Technical AI Safety Project Sprint (Mentee)
March 2026 - June 2026
Selected for hands-on mechanistic interpretability research under expert mentorship, probing for deceptive alignment in LLMs. Grew into my ongoing Mechanistic Audit of Deception Detection project.
BlueDot Impact — Technical AI Safety Course
February 2026 - April 2026
Intensive training on core safety topics: frontier-lab safety policies, emergent capabilities, model evaluations, and techniques for training safer systems.
IBM Research — DeepScanner Internship
June 2025 - September 2025
Co-designed and developed IBM DeepScanner, a mechanistic interpretability tool detecting anomalous patterns (hallucinations, toxicity) in LLM activations using non-parametric scan statistics.
Paper Replications
Ongoing
Replicated recent work on deceptive alignment and circuit analysis, including Apollo Research's probes, the LIAR bench, and Matryoshka Sparse Autoencoders (documented in a public LessWrong writeup).
Technical Literature Reviews
Ongoing
In-depth reviews of monosemanticity, probing classifiers, Sparse Autoencoders, Representation Engineering, and the trade-offs between mechanistic interpretability and other safety approaches.