Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling

Federico Tiblias, Irina Bigoulaeva, Jingcheng Niu, Simone Balloccu and Iryna Gurevych.
TMLR 2026

TL;DR

Supervised Multi-Dimensional Scaling (SMDS) is a model-agnostic method to automatically discover feature manifolds in language models, where prior efforts targeted specific geometries for specific features and thus lacked generalization. Applying SMDS to temporal reasoning as a case study, we find features forming circles, lines, and clusters whose geometry consistently reflects the properties of the concepts they represent, stays stable across model families and sizes, actively supports reasoning, and dynamically reshapes in response to context changes — supporting a model of entity-based reasoning in which LMs encode and transform structured representations.

How to Cite

@article{tiblias2026hypothesisdriven,
    title={Hypothesis-Driven Feature Manifold Analysis in {LLM}s via Supervised Multi-Dimensional Scaling},
    author={Federico Tiblias and Irina Bigoulaeva and Jingcheng Niu and Simone Balloccu and Iryna Gurevych},
    journal={Transactions on Machine Learning Research},
    issn={2835-8856},
    year={2026},
    url={https://openreview.net/forum?id=vCKZ40YYPr},
    note={}
}