EEG-PRISM: Physiologically-Grounded Interpretability of Predictions by EEG Foundation Models
EEG-PRISM is a post-hoc interpretability method that maps time-channel attribution scores from EEG foundation models into physiologically meaningful frequency and source domains without modifying or retraining the underlying model The method uses linear transformations and backpropagation rules, with mappings to the frequency domain via invertible DFT and to the source domain via an approximately invertible EEG generative model In simulation, EEG-PRISM achieves near-perfect spectral recovery and
Analysis
TL;DR
- EEG-PRISM is a post-hoc interpretability method that maps time-channel attribution scores from EEG foundation models into physiologically meaningful frequency and source domains without modifying or retraining the underlying model
- The method uses linear transformations and backpropagation rules, with mappings to the frequency domain via invertible DFT and to the source domain via an approximately invertible EEG generative model
- In simulation, EEG-PRISM achieves near-perfect spectral recovery and 69.2% spatial accuracy
- Applied to real clinical data, it correctly identifies delta-theta activity as most salient in epilepsy and localizes seizure onset regions with 50% accuracy
- In autism analysis, EEG-PRISM localizes predictive delta-alpha biomarkers to frontal and temporal regions, consistent with prior clinical findings
Why It Matters
EEG foundation models are rapidly advancing AI for brain signal analysis, but their interpretability has lagged behind—existing explainable AI techniques produce attribution scores in time-channel space, which clinicians find unintuitive and disconnected from established EEG knowledge. EEG-PRISM bridges this gap by providing a universal, model-agnostic interpretability layer that translates model attributions into domains clinicians already understand, enabling trust, validation, and clinical adoption of foundation model predictions.
Technical Details
- EEG-PRISM is a theoretically grounded post-hoc attribution method that leverages linear transformations and established backpropagation rules to remap attribution scores from the time-channel input space into alternative domains
- Frequency domain mapping is achieved via an invertible Discrete Fourier Transform (DFT), preserving attribution fidelity across spectral bands
- Source domain mapping uses an approximately invertible EEG generative model to project attributions from scalp-level electrodes into spatial source locations
- Evaluated across five foundation models and four AI explainers on both simulated and real clinical datasets (epilepsy and autism)
- Simulation results show near-perfect spectral recovery and 69.2% spatial accuracy; real-data evaluations demonstrate clinically consistent biomarker localization
Industry Insight
- The method's model-agnostic design means it can be applied to any existing EEG foundation model as a drop-in interpretability layer, making it immediately useful for practitioners without requiring model retraining or architectural changes
- As EEG foundation models move toward clinical deployment, physiologically-grounded interpretability will become a regulatory and trust prerequisite—EEG-PRISM provides a blueprint for how XAI methods should align with domain-specific clinical intuition
- The dual capability for both window-level transient event analysis (e.g., seizure localization) and group-level biomarker identification positions this approach as a versatile tool spanning diagnostic, monitoring, and research use cases in clinical neuroscience
Disclaimer: The above content is generated by AI and is for reference only.