{"id":"APP-260814-0000","repo_url":"https://github.com/lccqqqqq/sae-feature-nonlocality","versions":[{"v":1,"tag":"v1.0.0","commit":"01b5e62080e4f2e0bb522c1ebd8bd0c2d45df106","app_publication_id":"app-v1:sha256:e4d1298d35df9ec4064e7d22442a978e4b98647496b5a2dca6f386fbb82d9e06","release_url":"https://github.com/lccqqqqq/sae-feature-nonlocality/releases/tag/v1.0.0","listed_at":"2026-08-14","title":"Measuring Semantic Abstractness of SAE Features via Nonlocality","authors":[{"name":"Chuqiao Lin","affiliation":"Rudolf Peierls Centre for Theoretical Physics, University of Oxford"},{"name":"Shivaji L. Sondhi","affiliation":"Rudolf Peierls Centre for Theoretical Physics, University of Oxford"},{"name":"Xiao-Liang Qi","affiliation":"Leinweber Institute for Theoretical Physics, Stanford University"}],"domain":"mechanistic-interpretability","tags":["sparse-autoencoders","feature-abstractness","nonlocality","steering","jailbreak-audit"],"paper_summary":"Sparse autoencoders (SAEs) decompose language-model activations into features, but knowing *what* a feature responds to (its description) does not settle *at what level of abstraction* it operates — a token-matching feature and a genuinely contextual one can carry similar descriptions. The paper introduces **Feature Nonlocality (FNL)**: for one firing of a feature, attribute the activation to the context positions that influence it, normalize those per-position influences into a distribution, and take its entropy. A feature driven by a single token has near-zero FNL; a feature integrating a whole passage has high FNL. The paper shows FNL behaves as an abstractness measure should — it rises with network depth, is stable across text corpora once a reliability ceiling is accounted for, separates contextual from token-driven features, and predicts robustness of activation under meaning-preserving paraphrase — and then uses it in two applications: auditing a published SAE-based jailbreak mitigation (finding its features are mostly positional artifacts, not detectors of harmful intent) and selecting features for steering (steering high-FNL features improves benchmark accuracy where steering low-FNL ones does not, though gains are model-specific)."}]}