My background is in mathematics, philosophy, and machine learning. I am working on incorporating ideas from philosophy of science and measurement theory into mechanistic interpretability.

For example: What body of evidence warrants a specific type of mechanistic claim? What assumptions does it carry? How do you tell that two mechanisms are the same? Is the mechanism a subspace, a feature, an attention head? How does it appear in weight-space vs activation-space vs training dynamics? Does the mechanism transfer to novel contexts?

I have developed frameworks (Mechanistic Validity, Mechanistic Views) to formalize and address these questions; I am now applying them to audit existing works and ground my own work — informing new experiments, controls, and calibrations.

Background

Visiting Researcher — University of Edinburgh, School of Informatics (2026–present)

Machine Learning Engineer — Eluve Inc (2024–2026)

Machine Learning Engineer — Swarm Labs (2023–2024)

M.S. Computer Science — University of Massachusetts Amherst (2020–2022)

B.S. Mathematics, Philosophy — University of Massachusetts Amherst (2016–2020)

Preprints

  • Elliot Tower. Weight-Space Discovery of an Epistatic Circuit Invisible to Activation-Based Methods. 2026. [Preprint]

  • Elliot Tower. Epistatic Circuits: Exhaustive Interaction Decomposition of Transformer Circuits. 2026. [Preprint]

  • Elliot Tower. Mechanistic Validity: A Validity Theory, Research Methodology, and Evidence Standard for Mechanistic Claims. 2026. [Preprint] [Website]

  • Elliot Tower. Mechanistic Views: An Atlas of Hidden Commitments and a Realism Criterion for Mechanistic Claims. 2026. [Preprint] [Website]

Selected Publications

  • Rajarshi Das, Ameya Godbole, Ankita Naik, Elliot Tower, Manzil Zaheer, Hannaneh Hajishirzi, Robin Jia, Andrew McCallum. Knowledge Base Question Answering by Case-based Reasoning over Subgraphs. ICML 2022. [Proceedings]

Current Research

1. Mechanistic Validity (Research Program)

Auditing published claims in mechanistic interpretability with pre-registered experiments and formal validity criteria.

2. Factorized Circuits

Decomposing pretrained transformer weights into shared factor banks with sparse selectors to analyze information flow through the residual stream.