-
Causal Abstractions of Neural NetworksLLMs/Interpretability 2026. 6. 18. 08:46
(NeurIPS 2021)
neural network를 구조적으로 causal abstraction하기 때문에
접근 방향이 다르긴 하지만,
causal abstraction 개념을 formal하게 쉽게 설명해주고
전반적으로 process를 진행하는 방법을 알려준다.
1. Hypothesis formulation 2. Alighment search 3. verifying experimentally - interchange intervention
오래 된 논문이라도 무시하면 안된다.



















'LLMs > Interpretability' 카테고리의 다른 글
How to use and interpret activation patching (0) 2026.06.18 The Hydra Effect: Emergent Self-repair in LM Computations (0) 2026.06.18 Features (0) 2026.06.17 Causal Abstraction (0) 2026.06.17 Circuits (0) 2026.06.17