ABOUT ME

-

Today
-
Yesterday
-
Total
-
  • [on-going (until next week)] research question
    LLMs/Interpretability 2026. 6. 10. 16:33

    여러번 blog에 쓴 적이 있다. (in causal inference, mechanistic interpretability category)

    "LLM이 causal inference를 수행할 때, 내부의 DAG structure를 찾을 수 있을까?" 라는 의문을 말이다.


    박사를 하게 된다면 해보고 싶은 연구였다..

     

    그런데 이제는 마음을 접었는데..

    이유인 즉슨,

    나를 받아주실 교수님이 계실지도 의문이고, 박사 입학이 합격 가능할지도 의문이기 때문이다..

    자신감과 용기를 잃었음..

    마음을 비우기로 하였다.

     

    아무튼, 그런데 이걸 왜 다시 꺼내들었냐면, 

    DGM 결석처리 관련하여, 개인 과제를 제출해야하는데, 

    기왕이면 해보고 싶었던 걸 좀 더 구체화해보고 싶기 때문이다.

     

    근데 이렇게 내가 하고 싶은 걸 해도 되는건지 잘 모르겠다.. ㅎㅎ 

     

    교수님은 정말 천사시다 ㅠㅠ 엉엉 

    내 인생 최초의 F를 면할 기회를 주셨다..... ㅠㅠ 


    아무튼 그래서 좀 구체화해보고자 하는데,

    찾아보니, 내가 가졌던 생각, 의문들 관련 최근 연구가 꽤 있다! 

    역시 나만 이런 생각을 하는 게 아녔어! -_-


    Do Language Models Implement Structural Causal Models? A Mechanistic Interpretability Study of Causal Reasoning.

    - Do LLMs' internal representations form a faithful causal abstraction of a causal model relevant to the task? 

    - Mechanistically test whether LLMs implement structural causal models internally.

    * The core research question:

    - Given a task with a known causal graph, can we identify internal LLM features corresponding to causal variables, and do interventions on those features produce downstream changes predicted by the graph?

    * Related literatures:

    causal inference, causal representation learning, and mechanistic interpretability

    * Existing works on "Mechanistic interpretability can discover causally implicated subnetworks of interpretable features"

    1) Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

    https://openreview.net/forum?id=I4e82CIDxv

    - introducing methods for discovering and editing feature-level causal circuits in language models and applies them to generalization failures caused by spurious corrlelations. 

     

    2)  Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

    https://www.jmlr.org/papers/volume26/23-0058/23-0058.pdf

    - proposing theoretical foundation for mechanistic interpretability, unifying activation patching, causal mediation analysis, causal tracing, sparse autoencoders, distributed alignment search, and steering under one causal framework.

     

    3) Interpretability at Scale: Identifying Causal Mechanisms in Alpaca

    https://papers.nips.cc/paper_files/paper/2023/hash/f6a8b109d4d4fd64c75e94aaf85d9697-Abstract-Conference.html

    - Distributed Alignment Search and Boundless DAS search for alignments between high-level causal variables and distributed neural representations. In Alpaca-7B, Boundless DAS found that a simple numerical reasoning behaviour could be explained by two interpretable Boolean variables aligned with internal representations.

     

    4) Causal Reasoning and Large Language Models: Opening a New Frontier for Causality

    https://openreview.net/forum?id=mqoxLkX210

    - reporting that LLMs can generate correct causal arguments across many tasks and may help experts set up causal analyses, but also emphasizing unpredictable failure modes and suggesting combining LLMs with existing casual techniques.

     

    5) Can Large Language Models Infer Causation from Correlation?

    https://arxiv.org/pdf/2306.05836

    - found that models were close to random on a benchmark designed to test causation-from-correlation reasoning and that fine-tuned models struggled to generalize out of distribution.

     

    6) CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs

    https://arxiv.org/pdf/2404.06349

    - reporting that LLMs struggle on large causal networks and collider structures, even when they do better on simpler chain-like structures.

     

    7) Causal Inference with Large Language Model: A Survey

    https://aclanthology.org/2025.findings-naacl.327.pdf

    * Research gap:

    Most exsting work uses causality to explain model behaviour. Our contribution could be to use mechanistic interpretability to study causal inference itself: whether the model internally represents variables, interventions, confounders, colliders, mediators, and counterfactuals in a way that supports causal reasoning.

     


     

Designed by Tistory.