Overview
Research investigated the relationship between brain-AI alignment and the causal involvement of Large Language Model (LLM) attention heads in computational processes. Specifically, it examined whether attention heads aligned with human electroencephalography (EEG) data during an abstract pattern-completion task were also critical for the LLM's task performance. The study employed an abstract pattern-completion task (AAABAAA $\rightarrow$ B) to compare LLM attention-head representations with human EEG signals and assessed the impact of ablating these heads on task performance.
Research Context
Brain-AI alignment is frequently interpreted as an indicator that models and brains execute similar computations. However, the causal role of these aligned units in a model's computation is seldom verified. This research aimed to address this gap by scrutinizing the causal involvement of brain-aligned attention heads in LLM operations, particularly on an abstract pattern-completion task. Prior interpretability work defines specific head sets, such as concept vectors (CVs) and function vectors (FVs), without reference to brain activity. CVs represent abstract patterns across different formats, while FVs are selected based on their contribution to correct-answer prediction. This study sought to integrate these existing interpretability frameworks with brain alignment data.
Approach
The study used an abstract pattern-completion task, specifically AAABAAA $\rightarrow$ B, to evaluate LLM behavior and human brain activity. LLM attention-head representations were compared with human EEG data. To assess causal involvement, the researchers performed ablation experiments on specific attention heads and measured the effect on task performance. Two main sets of attention heads were compared:
- Heads aligned with human EEG data (brain-aligned heads).
- Heads selected via attribution patching, a method not referencing brain activity.
- Concept vectors (CVs), which represent abstract patterns across formats.
- Function vectors (FVs), chosen for their contribution to correct-answer prediction.
The study spanned 17 LLM models, ranging in size from 3 billion to 72 billion parameters.
Findings
The research revealed a dissociation between brain alignment and causation in LLM attention heads:
- Brain-aligned heads were observed to contribute to task performance.
- However, the removal of brain-aligned heads was substantially less disruptive to performance compared to the removal of heads selected via attribution patching. This suggests that while brain-aligned heads participate, they are not as critically causal for task success as attribution-selected heads.
- Brain alignment showed minimal association with Function Vector (FV) scores, indicating that heads critical for predicting the correct answer are not necessarily those aligned with human brain activity.
- The association between brain alignment and Concept Vector (CV) scores varied across the different LLM models tested.
- Among brain-aligned heads, two recurring attention profiles were identified:
- Novelty heads: These emphasize distinctive elements and track salience, attending to the same elements that humans focus on. Despite this, their removal resulted in less damage to task performance on average than random ablation.
- Repetition heads: These focus on repeating elements and contribute modestly to performance. They are associated with the representation of abstract patterns (CVs).
- Across the 17 models evaluated (ranging from 3B to 72B parameters), ablating heads ranked by FV scores (FV-ranked removal) was substantially more disruptive to performance than ablating heads ranked by brain alignment (brain-ranked removal).
- Overall, brain alignment primarily captured how the model processes or 'reads' the stimulus, rather than how it represents the underlying pattern or arrives at the task solution.
Why This Matters
This research provides a nuanced understanding of brain-AI alignment, indicating that mere alignment does not directly equate to causal involvement in a model's core computational mechanisms for task solving. It suggests that metrics beyond brain alignment may be necessary to identify the most causally significant components within LLMs. The findings highlight that brain alignment might reflect peripheral processing like stimulus parsing rather than the central abstract pattern recognition or solution generation pathways.