Skip to Main Content

How I am building my foundations in AI safety research — through mentored research, structured training, and hands-on replications.

Research Assistant — Jinesis AI Lab, University of Toronto

June 2026 - Present

Researching the generalizability of mechanistic interpretability techniques: when do results established on one model transfer to other models, tasks, and settings?

BlueDot Impact — Technical AI Safety Project Sprint (Mentee)

March 2026 - June 2026

Selected for hands-on mechanistic interpretability research under expert mentorship, probing for deceptive alignment in LLMs. Grew into my ongoing Mechanistic Audit of Deception Detection project.

BlueDot Impact — Technical AI Safety Course

February 2026 - April 2026

Intensive training on core safety topics: frontier-lab safety policies, emergent capabilities, model evaluations, and techniques for training safer systems.

IBM Research — DeepScanner Internship

June 2025 - September 2025

Co-designed and developed IBM DeepScanner, a mechanistic interpretability tool detecting anomalous patterns (hallucinations, toxicity) in LLM activations using non-parametric scan statistics.

Paper Replications

Ongoing

Replicated recent work on deceptive alignment and circuit analysis, including Apollo Research's probes, the LIAR bench, and Matryoshka Sparse Autoencoders (documented in a public LessWrong writeup).

Technical Literature Reviews

Ongoing

In-depth reviews of monosemanticity, probing classifiers, Sparse Autoencoders, Representation Engineering, and the trade-offs between mechanistic interpretability and other safety approaches.