-
(1/3) A Primer on the Inner Workings of Transformer-based LMsLLMs/Interpretability 2026. 6. 12. 15:47
https://arxiv.org/pdf/2405.00208v3
(Oct 2024)


Figure 1: Survey overview. Section 2 introduces the Transformer language model and its components. Section 3 and Section 4 present interpretability techniques used to analyze models’ inner workings. Finally, Section 5 presents known inner workings of Transformer language models.






Figure 2: Unrolled Transformer LM with expanded views of the Attention and Feedforward network blocks, including model weights (gray) and residual stream states (green). Based on figures from (Ferrando & Voita, 2024; Voita et al., 2023) .



Figure 3: Forward pass decomposition in a simplified Transformer LM. The direct path (red), full OV circuits (yellow) and virtual attention heads (grey) expressed in Equation 11 are highlighted.




















'LLMs > Interpretability' 카테고리의 다른 글
(3/3) A Primer on the Inner Workings of Transformer-based LMs (0) 2026.06.13 (2/3) A Primer on the Inner Workings of Transformer-based LMs (0) 2026.06.13 [흥미로운 예시] Automating psychological hypothesis generation with AI: when LLMs meet causal graph (0) 2026.06.11 (Review) Mechanistic Interpretability (0) 2026.06.11 [Draft] Testing SCM Faithfulness in LLM Internals (0) 2026.06.10