Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
Self-harm information in language models is concentrated in the final 3-7% of network layers (93-97% depth), as revealed by linear probe experiments across two datasets (X-Sensitive and SH-Detection). The most accurate linear probes for self-harm detection are not necessarily the most linearly separable, indicating a complex representation of self-harm content. Gemma-3-4B represents the contrastive self-harm direction in a more intricate way compared to other models, suggesting architectural dif
Analysis
TL;DR
- Self-harm information in language models is concentrated in the final 3-7% of network layers (93-97% depth), as revealed by linear probe experiments across two datasets (X-Sensitive and SH-Detection).
- The most accurate linear probes for self-harm detection are not necessarily the most linearly separable, indicating a complex representation of self-harm content.
- Gemma-3-4B represents the contrastive self-harm direction in a more intricate way compared to other models, suggesting architectural differences in how self-harm is encoded.
Why It Matters
This research is critical for improving self-harm detection systems in AI, as it highlights the importance of focusing on deeper layers of language models for accurate identification. The findings also emphasize the need for tailored interventions and governance strategies to address self-harm content effectively, particularly in high-stakes applications like user safety and mental health support.
Technical Details
- Linear probes were trained across all layers of four language models on two self-harm datasets (X-Sensitive and SH-Detection) to analyze where self-harm information is encoded.
- The study found that self-harm signals are most prominent in the final 3-7% of network layers, indicating that deeper layers are crucial for capturing nuanced self-harm representations.
- Contrastive self-harm directions were extracted and normalized, revealing that the most accurate probes are not always the most linearly separable, suggesting a non-linear relationship between self-harm content and model representations.
- Gemma-3-4B exhibited a more intricate representation of the contrastive self-harm direction compared to other models, highlighting architectural differences in how self-harm is processed.
Industry Insight
- AI developers should prioritize optimizing the deeper layers of language models for self-harm detection tasks, as these layers contain the most relevant information for accurate identification.
- The non-linear relationship between self-harm representations and probe accuracy suggests that traditional linear methods may not be sufficient, prompting the exploration of more advanced techniques for self-harm detection.
- The unique representation of self-harm in Gemma-3-4B indicates that model architecture plays a significant role in how sensitive content is encoded, which should be considered when selecting or designing models for safety-critical applications.
Disclaimer: The above content is generated by AI and is for reference only.