- Rudolf Peierls Centre for Theoretical Physics, University of Oxford
- Leinweber Institute for Theoretical Physics, Stanford University
APP-260814-0000mechanistic-interpretabilityRelease v1.0.0Listed Aug 14, 2026
Open in your agent
git clone --recurse-submodules --branch v1.0.0 https://github.com/lccqqqqq/sae-feature-nonlocality.gitcd sae-feature-nonlocalityclaudegit clone --recurse-submodules --branch v1.0.0 https://github.com/lccqqqqq/sae-feature-nonlocality.gitcd sae-feature-nonlocalitycodexThese commands download the paper and start the agent in its folder. Then ask it anything about the work, or ask it to reproduce a figure. Other coding agents work too.
Working in another project? With the APP plugin, run /load-paper https://github.com/lccqqqqq/sae-feature-nonlocality/releases/tag/v1.0.0 to bring this paper into it.
Summary
Sparse autoencoders (SAEs) decompose language-model activations into features, but knowing what a feature responds to (its description) does not settle at what level of abstraction it operates — a token-matching feature and a genuinely contextual one can carry similar descriptions. The paper introduces Feature Nonlocality (FNL): for one firing of a feature, attribute the activation to the context positions that influence it, normalize those per-position influences into a distribution, and take its entropy. A feature driven by a single token has near-zero FNL; a feature integrating a whole passage has high FNL. The paper shows FNL behaves as an abstractness measure should — it rises with network depth, is stable across text corpora once a reliability ceiling is accounted for, separates contextual from token-driven features, and predicts robustness of activation under meaning-preserving paraphrase — and then uses it in two applications: auditing a published SAE-based jailbreak mitigation (finding its features are mostly positional artifacts, not detectors of harmful intent) and selecting features for steering (steering high-FNL features improves benchmark accuracy where steering low-FNL ones does not, though gains are model-specific).