Publication Highlights

Alignment & Interpretability

Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families
Separates what refusal reads from what models understand, showing refusal often routes through a narrow harm-keyed channel rather than the broader moral subspace.
arXiv preprint, 2026
Calibrating Interpretability Instruments Before Trusting Their Verdicts
Documents failure modes in causal interpretability measurements and gives calibration checks before treating instrument readings as findings.
arXiv preprint, 2026
How Language Models Organize and Structure Moral Knowledge
Maps how language models represent moral foundations geometrically, finding distinct but integrated directions that encode moral tension rather than a pre-resolved verdict.
arXiv preprint, 2026
Output Dilution: Redundant but Fragile Representations in MoE Models
Shows that MoE models can encode moral content redundantly while remaining fragile because expert outputs are diluted before reaching the residual stream.
arXiv preprint, 2026
When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis
Introduces activation-noise fragility as a companion to probe accuracy for tracking representation robustness across pretraining.
arXiv preprint, 2026

Interpretability & Infrastructure

ExecuTorch: A Unified PyTorch Solution to Run AI Models On-Device
Mergen Nachin, Digant Desai, Sicheng Stephen Jia, Chen Lai, Mengwei Liu, Jacob Szwejbka, Raziel Alvarez, RJ Ascani, Dave Bort, Manuel Candales, Andrew Caples, Yanan Cao, Zhengxu Chen, Soumith Chintala, Gregory Comer, Tanvir Islam, Songhao Jia, Tarun Karuturi, Jack Khuu, Abhinay Kukkadapu, Tugsbayasgalan Manlaibaatar, Andrew Or, Kimish Patel, Siddartha Pothapragada, Lucy Qiu, Supriya Rao, Orion Reblitz-Richardson, Max Ren, Scott Roy, Anthony Shoumikhin, Scott Wolchok, Guang Yang, Angela Yi, Martin Yuan, Hansong Zhang, Jack Zhang, Jerry Zhang, Shunting Zhang, C. Cagatay Bilgin
arXiv preprint, 2026
Captum: A unified and generic model interpretability library for PyTorch
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, Orion Reblitz-Richardson
arXiv preprint, 2020
Investigating Saturation Effects in Integrated Gradients
Vivek Miglani, Narine Kokhlikyan, Bilal Alsallakh, Orion Reblitz-Richardson
Preprint, 2020