-
How to use and interpret activation patchingLLMs/Interpretability 2026. 6. 18. 21:14
(Apr 2024)
Activation patching을 어떻게 할 것인지에 대한 구체적 설명

















'LLMs > Interpretability' 카테고리의 다른 글
A Circuit for Indirect Object Identification in GPT-2 Small (0) 2026.06.19 Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations (0) 2026.06.18 The Hydra Effect: Emergent Self-repair in LM Computations (0) 2026.06.18 Causal Abstractions of Neural Networks (0) 2026.06.18 Features (0) 2026.06.17