Skip to content

Measuring Semantic Abstractness of SAE Features via Nonlocality

Chuqiao Lin1, Shivaji L. Sondhi1, Xiao-Liang Qi2

  1. Rudolf Peierls Centre for Theoretical Physics, University of Oxford
  2. Leinweber Institute for Theoretical Physics, Stanford University

APP-260814-0000mechanistic-interpretabilityRelease v1.0.0Listed Aug 14, 2026

Open in your agent

Terminal window
git clone --recurse-submodules --branch v1.0.0 https://github.com/lccqqqqq/sae-feature-nonlocality.git
cd sae-feature-nonlocality
claude

These commands download the paper and start the agent in its folder. Then ask it anything about the work, or ask it to reproduce a figure. Other coding agents work too.

Working in another project? With the APP plugin, run /load-paper https://github.com/lccqqqqq/sae-feature-nonlocality/releases/tag/v1.0.0 to bring this paper into it.

View the repository on GitHub

Summary

Sparse autoencoders (SAEs) decompose language-model activations into features, but knowing what a feature responds to (its description) does not settle at what level of abstraction it operates — a token-matching feature and a genuinely contextual one can carry similar descriptions. The paper introduces Feature Nonlocality (FNL): for one firing of a feature, attribute the activation to the context positions that influence it, normalize those per-position influences into a distribution, and take its entropy. A feature driven by a single token has near-zero FNL; a feature integrating a whole passage has high FNL. The paper shows FNL behaves as an abstractness measure should — it rises with network depth, is stable across text corpora once a reliability ceiling is accounted for, separates contextual from token-driven features, and predicts robustness of activation under meaning-preserving paraphrase — and then uses it in two applications: auditing a published SAE-based jailbreak mitigation (finding its features are mostly positional artifacts, not detectors of harmful intent) and selecting features for steering (steering high-FNL features improves benchmark accuracy where steering low-FNL ones does not, though gains are model-specific).

  • sparse-autoencoders
  • feature-abstractness
  • nonlocality
  • steering
  • jailbreak-audit

Protocol on GitHubRegistry on GitHubPaper agents on Zulip