-
Causal AbstractionLLMs/Interpretability 2026. 6. 17. 09:45
Finding interpretable high-level causal abstractions in lower-level neural networks. These methods involve a computationally expensive search and assume high-level variables align with groups of units or neurons.
To overcome the limitations, Geiger et al. (2023b) propose distributed alignment search (DAS), which performs distributed interchange interventions (DII) on non-basis-aligned subspaces of the low-level representation space found via gradient descent.
DAS interventions have been shown to be effective in finding features with causal influence in targeted syntactic evaluation (Arora et al., 2024), and in isolating the causal effect of individual attributes of entities (Huang et al. 2024a).
A DAS variant named Boundless DAS has been used to search for interpretable causal structure in large language models (Wu et al., 2023b). In this context, Causal Proxy Models (CPMs) were proposed as interpretable proxies trained to mimic the predictions of lower-level models and simulate their counterfactual behavior after targeted interventions (Wu et al., 2023a).
Causal Abstractions of Neural Networks
Inducing causal structure for interpretable neural networks
https://arxiv.org/pdf/2112.00826
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
https://arxiv.org/pdf/2301.04709
Finding alignments between interpretable causal variables and distributed neural representations
https://arxiv.org/pdf/2303.02536
Interpretability at scale: Identifying causal mechanisms in alpaca
Causal proxy models for concept-based model explanations
https://arxiv.org/pdf/2209.14279
'LLMs > Interpretability' 카테고리의 다른 글
Causal Abstractions of Neural Networks (0) 2026.06.18 Features (0) 2026.06.17 Circuits (0) 2026.06.17 Activation Patching (0) 2026.06.17 생성모델 관점에서 생각해보더라도 (0) 2026.06.15