-
생성모델 관점에서 생각해보더라도LLMs/Interpretability 2026. 6. 15. 12:26
latent space를 갖는 generative model의 관점에서 생각해본다면,
latent space에서 causally disentangled variables (variation of factors)를 찾는다면, data generating process (mechanism)을 이해할 수 있잖아.
마찬가지로 Language model도 causal representation (feature)와 causal graph (circuit)을 통해서 추론하고 있다는 사실을 MI이 보인다는 건, LM이 data generating process를 이해하고 있다는 (혹은 이해할 수 있다는) 사실이 아닐까?
나는 MI가 LLM의 이러한 능력을 근거있게 증명해줄 수 있는 중요한 tool이라고 생각돼.
PFN에 흠뻑 빠질 수 밖에 없었던 건, 정말 세상의 모든 SCM을 다 배운다면, 전지전능한 agent가 될 수 있지 않을까? 하는 생각이 들었었고,
그것보다는 좀 더 현실적으로,
Robust Agents Learn Causal World Models 논문에서 보였듯이,
https://letter-night.tistory.com/884
Robust Agents Learn Causal World Models
충격 그 잡채.시사하는 바가 많아서 한번에 정리가 안된다. 이 논문은.. 그간 내가 공부를 하며 쌓아온 의문점들에 대해서 수식적인 증명을 통해 명확하게 답변한다. 천천히 다시 읽고 정리해야
letter-night.tistory.com
Causal Beysian Network를 학습함으로써 새로운 환경에서의 추론, 새로운 과학적 사실, 인과관계 규명을 할 수 있지 않을까?
하는 생각을 해본다.
난 결국 나의 이 Big question을 풀지 못한 채 졸업하고, 생계전선으로 향하는 것인가? ㅋㅋㅋ

'LLMs > Interpretability' 카테고리의 다른 글
Circuits (0) 2026.06.17 Activation Patching (0) 2026.06.17 어떻게 보면, LLM은 거대한 graph database retrieval system 같기도 해 (0) 2026.06.15 (2/2) Right Mediator: MI Through the Lens of Causal Mediation Analysis (0) 2026.06.14 (1/2) Right Mediator: MI Through the Lens of Causal Mediation Analysis (0) 2026.06.14