-
Activation PatchingLLMs/Interpretability 2026. 6. 17. 08:24
Activation Patching (Interchange Interventions)
We can view the computations of a Transformer-based LM as a causal model, and use causality tools to shed light on the contribution to the prediction of each model component across different positions.
Causal Abstractions of Neural Networks
The Hydra Effect: Emergent Self-repair in Language Model Computations
https://arxiv.org/pdf/2307.15771
How to use and interpret activation patching
https://arxiv.org/pdf/2404.15255
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
https://arxiv.org/pdf/2303.02536
'LLMs > Interpretability' 카테고리의 다른 글
Causal Abstraction (0) 2026.06.17 Circuits (0) 2026.06.17 생성모델 관점에서 생각해보더라도 (0) 2026.06.15 어떻게 보면, LLM은 거대한 graph database retrieval system 같기도 해 (0) 2026.06.15 (2/2) Right Mediator: MI Through the Lens of Causal Mediation Analysis (0) 2026.06.14