-
The Hydra Effect: Emergent Self-repair in LM ComputationsLLMs/Interpretability 2026. 6. 18. 11:16
(Jul 2023 DeepMind)
https://arxiv.org/pdf/2307.15771
MI로 LM의 internal computation 에서 hydra effect가 존재함을 밝힌다.
본 논문에서 밝혀낸 사실도 흥미롭긴 하지만,
본 논문을 본 이유는,
LM을 causal model로 보고, computation graph를 causal graph로 해석해서, intervention을 통해 effect를 estimate하는 방법 (activation patching)을 어떻게 이론적으로 정립하고, 실험적으로 보였는지를 살펴보기 위함이다.
내가 가진 RQ과 관련하여 중요한 부분은 아래 내용이다.
neural network의 내부 구조를 SCM으로 보고, causality를 tool로 사용해서 내부 작용을 분석하지만,
causal reasoning을 하는가에 대한 분석은 아니다.
내가 하고자 하는 건, 바로 이 causal reasoning에 대한 것이다.



















'LLMs > Interpretability' 카테고리의 다른 글
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations (0) 2026.06.18 How to use and interpret activation patching (0) 2026.06.18 Causal Abstractions of Neural Networks (0) 2026.06.18 Features (0) 2026.06.17 Causal Abstraction (0) 2026.06.17