-
(Review) Mechanistic InterpretabilityLLMs/Interpretability 2026. 6. 11. 07:36
Mechanistic Interpretability literature 를 공부하는 건 만만치 않다.
전반적인 landscape를 파악하고, 반드시 알아야 하는 수학적 tool들은 아래와 같을 것 같다.
금새 뚝닥 볼 수 있는 내용이 아니다. 아마 이번 주말 내내 보고 있을 듯 ㅎㅎ
1. Circuits
https://distill.pub/2020/circuits/zoom-in/
Zoom In: An Introduction to Circuits
By studying the connections between neurons, we can find meaningful algorithms in the weights of neural networks.
distill.pub
2. Transformer Circuits
https://transformer-circuits.pub/2021/framework/index.html
A Mathematical Framework for Transformer Circuits
Contents Transformer language models are an emerging technology that is gaining increasingly broad real-world use, for example in systems like GPT-3 , LaMDA , Codex , Meena , Gopher , and similar models. However, as these models scale, their open-endedne
transformer-circuits.pub
3. Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
https://arxiv.org/pdf/2408.01416
4. A Primer on the Inner Workings of Transformer-based Language Model
https://arxiv.org/html/2405.00208v3
A Primer on the Inner Workings of Transformer-based Language Models
A Primer on the Inner Workings of Transformer-based Language Models Javier Ferrando1 , Gabriele Sarti2, Arianna Bisazza2, Marta R. Costa-jussà3 1Universitat Politècnica de Catalunya, 2CLCG, University of Groningen, 3FAIR, Meta Correspondence to: jferran
arxiv.org
5. Open Problems in Mechanistic Interpretability
https://arxiv.org/pdf/2501.16496
'LLMs > Interpretability' 카테고리의 다른 글