Research · 4 min read

The Effect Cloud: FEGA Shows Sparse Autoencoder Features Rarely Steer Like Simple Directions

A July 27 arXiv paper introduces FEGA, showing that sparse autoencoder features split into value-like and pointer-like types — and that interpretable features rarely steer models along stable directions.

By Classy AI News · July 29, 2026

The Effect Cloud: FEGA Shows Sparse Autoencoder Features Rarely Steer Like Simple Directions

Sparse autoencoders have become the interpretability tool of choice for researchers trying to map neural network activations onto human-readable features. Yet a persistent frustration keeps surfacing in the literature: a feature that activates cleanly on "Golden Gate Bridge" may steer the model unpredictably, weakly, or in the wrong direction when you try to use it for control.

A paper posted to arXiv on July 27, 2026, offers a geometric explanation for that mismatch — and it may reshape how the field thinks about SAE-based steering.

The problem FEGA sets out to solve

The work, titled "Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects," introduces Feature-Effect Geometry Analysis (FEGA). Authors Phu Gia Hoang and colleagues study not where features live inside the model, but what happens to model logits when you remove an active SAE feature across many contexts.

The abstract states the core finding plainly: "Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change."

Prior interpretability work focused on feature geometry inside the network — how activations cluster in representation space. FEGA inverts the question: what is the geometry of the output changes those features cause?

3D human brain in a digital environment evoking neural representation geometry

Value-like versus pointer-like features

Across multiple SAE variants, FEGA finds that consistent one-dimensional effects are rare. Few features behave like reusable steering directions you can apply uniformly across prompts.

To make sense of the variation, the authors distinguish two categories:

Value-like features tie to static information — factual attributes, entity properties, domain knowledge that does not depend heavily on conversational context. These more often exhibit structured, low-dimensional logit effects, though even they typically span several directions rather than one clean axis.

Pointer-like features associate with context-dependent operations — routing attention, selecting among possible continuations, triggering procedural behaviors. These predominantly produce diffuse, high-dimensional logit changes that resist simple steering.

The distinction matters practically. Interpretability researchers have spent years labeling SAE features from activation patterns and assuming those labels predict causal behavior. FEGA shows a feature can be both interpretable and causally relevant while still failing to provide a stable steering handle.

Why steering fails even when labels look right

Consider a feature that reliably activates on medical terminology. Under activation-based selection, you might expect ablating it to reduce medical content in outputs. FEGA's framework predicts the outcome depends on whether the feature is value-like (storing medical knowledge) or pointer-like (routing to medical response modes depending on prompt structure).

Pointer-like features explain why steering experiments often show prompt-dependent effects — sometimes amplifying the intended behavior, sometimes suppressing it, sometimes doing nothing measurable. The feature is doing real computational work; it just is not doing work that compresses to a single logit direction.

This aligns with growing empirical reports from labs running SAE steering at scale. Features that look crisp under the microscope behave messily under the scalpel.

Abstract AI network structure representing downstream logit effects

Implications for mechanistic interpretability

If FEGA's taxonomy holds across model families, several research programs need recalibration.

Feature labeling pipelines should treat activation descriptions as necessary but insufficient evidence of causal role. Downstream effect geometry needs to be measured, not inferred.

Steering benchmarks should report success rates separately for value-like and pointer-like features. Aggregate steering accuracy numbers may hide a bimodal distribution: reliable control on static-knowledge features, unreliable control on operational features.

Safety applications that depend on SAE-based monitoring — detecting deception features, suppressing harmful concept directions — must account for the possibility that the most safety-relevant features are pointer-like and therefore hardest to steer predictably.

What the paper does not claim

The arXiv preprint (2607.24645) is an empirical and geometric analysis, not a new SAE architecture. It does not propose a fix for steering instability. It does not establish causal mechanisms for why pointer-like features produce diffuse effects — only that they do, consistently across the SAE variants tested.

Peer review and replication across larger models will determine how broadly the value/pointer taxonomy applies. For now, the paper gives interpretability researchers a vocabulary for a failure mode they have been encountering without a unified framework to describe it.

The bottom line

Mechanistic interpretability promised that understanding internal features would unlock control. FEGA suggests the link is more conditional than the field hoped: interpretable features encode both stored concepts and dynamic functions, and only the former reliably steer like directions in logit space.

For anyone building safety tools on top of sparse autoencoders, that is not a minor footnote. It is a design constraint.

Sources

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.