Publication Highlights
Alignment & Interpretability
Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families
Separates what refusal reads from what models understand, showing refusal often routes through a narrow harm-keyed channel rather than the broader moral subspace.
arXiv preprint, 2026
Calibrating Interpretability Instruments Before Trusting Their Verdicts
Documents failure modes in causal interpretability measurements and gives calibration checks before treating instrument readings as findings.
arXiv preprint, 2026
How Language Models Organize and Structure Moral Knowledge
Maps how language models represent moral foundations geometrically, finding distinct but integrated directions that encode moral tension rather than a pre-resolved verdict.
arXiv preprint, 2026
Output Dilution: Redundant but Fragile Representations in MoE Models
Shows that MoE models can encode moral content redundantly while remaining fragile because expert outputs are diluted before reaching the residual stream.
arXiv preprint, 2026
When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis
Introduces activation-noise fragility as a companion to probe accuracy for tracking representation robustness across pretraining.
arXiv preprint, 2026
Interpretability & Infrastructure
Captum: A unified and generic model interpretability library for PyTorch
arXiv preprint, 2020