No country for old linguists: LLM-brain alignment underdetermines neural computation
LLM-brain representational alignment can constrain mechanistic hypotheses but cannot by itself identify the underlying neural mechanism Murphy argues that Nastase et al. (2026) overreach by moving from alignment evidence to claims of "shared computational principles" and "fully mechanistic models" of language Three forms of underdetermination are identified: logical, causal, and computational — each undermining the leap from correlation to mechanistic explanation Encoding models can capture feat
Analysis
TL;DR
- LLM-brain representational alignment can constrain mechanistic hypotheses but cannot by itself identify the underlying neural mechanism
- Murphy argues that Nastase et al. (2026) overreach by moving from alignment evidence to claims of "shared computational principles" and "fully mechanistic models" of language
- Three forms of underdetermination are identified: logical, causal, and computational — each undermining the leap from correlation to mechanistic explanation
- Encoding models can capture features represented in neural activity without establishing that LLMs and biological brains share architecture or algorithm
Why It Matters
This paper raises a critical methodological concern for the rapidly growing field of LLM-brain alignment research: high representational similarity between models and neural data does not justify strong mechanistic claims. For AI practitioners and cognitive scientists alike, it serves as a cautionary note against overinterpreting alignment results as evidence of shared computation.
Technical Details
- The paper critiques the inferential gap between representational alignment (measured via encoding models) and mechanistic equivalence, emphasizing that multiple distinct architectures can produce similar representational patterns
- Murphy identifies three specific underdetermination problems: logical (alignment is consistent with multiple mechanistic hypotheses), causal (correlation does not establish that LLMs causally mirror brain computation), and computational (different algorithms can yield functionally equivalent representations)
- The analysis engages with Nastase et al.'s (2026) claim that LLMs instantiate the same computational principles as biological brains and can serve as "fully mechanistic models" of natural language processing
- The paper operates within the intersection of computational linguistics and computational neuroscience, addressing the epistemic status of alignment-based inference
Industry Insight
- Researchers should treat LLM-brain alignment metrics as hypothesis-generating rather than hypothesis-confirming; strong mechanistic claims require independent causal or architectural evidence
- The field would benefit from developing stricter inferential standards that distinguish representational similarity from mechanistic equivalence
- Practitioners building neuro-inspired AI systems should be cautious about assuming that alignment performance directly translates to biological plausibility or mechanistic insight
Disclaimer: The above content is generated by AI and is for reference only.