Research · 2 min read

LLM Information Geometry Is Shared Across Architectures and Can Steer Safely

A 10 September 2026 arXiv paper reports that large language models share Fisher Rao output geometry across transformer, state space, and recurrent architectures, enabling minimum disturbance control for steering and fine tuning.

By Classy AI News · September 14, 2026

LLM Information Geometry Is Shared Across Architectures and Can Steer Safely

What changed

Researchers posted arXiv:2609.11063 on 10 September 2026, titled The information geometry of large language models is shared, learned, and controllable. The work studies the Fisher Rao geometry of next token probability distributions rather than hidden activations, arguing that behavior is determined up to output preserving symmetries.

Across transformer, state space, and recurrent models, the authors report that output geometries agree more strongly than activation geometries. They show shared geometry supports semantic category transfer, that agreement with human word choices rises with scale and training, and that the geometry prescribes minimum disturbance local interventions for steering, editing, attribution, dictionary learning, and fine tuning.

Randomized experiments in the paper also report that deeper evidence substantially delays factual acquisition across every tested architecture and evidence construction.

Neural network visualization on a computer monitor

Why it matters

Teams building alignment, red teaming, and enterprise guardrails often fight activation level hooks that break when a vendor swaps layer norms or context length. If output geometry is the stable object, safety and steering tools could transfer across model families with less re tuning.

The paper’s control experiments claim geometric corrections improve steering while better preserving behavior on reference prompts than Euclidean control, and that updates learned on donor prompts transfer to unseen prompts. For product leaders, that is a path toward reusable safety patches rather than one off fine tunes per release.

The evidence depth result is equally practical: models may memorize surface patterns before absorbing deeper proof. Evaluation suites that only test shallow retrieval could green light models that fail under adversarial depth.

Who is affected

Applied ML teams shipping steering, RAG guardrails, or domain adapters should read the geometry based control section before committing to another LoRA stack.

Safety and eval vendors can benchmark whether their monitors track output geometry invariants instead of brittle activation probes.

Enterprise buyers comparing open and closed models gain a testable claim: architectures may differ internally while sharing external behavior geometry.

What to do next

Replicate one minimum disturbance steering experiment on a model you deploy. If geometric control preserves reference behavior better than your current hook, document the delta in your model card and procurement checklist.

Data science team reviewing model metrics on screens

What to watch

Whether any frontier lab publishes held out replication on the evidence depth delay claim using production scale corpora, not paper scale setups. Also watch if steering vendors adopt Fisher Rao metrics as a standard acceptance test in enterprise contracts.

Sources

  1. Primary. arXiv, The information geometry of large language models is shared, learned, and controllable (10 September 2026). Full methods, architecture comparisons, and control experiments.
  1. Secondary. arXiv cs.LG listing, Machine Learning September 2026 (September 2026). Confirms submission date and categorization.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.