More Data Cannot Break a Symmetry: Identifiability by Design
Unsupervised representational alignment is fundamentally bounded by the automorphism group of stimulus geometry, meaning more data alone cannot resolve certain symmetries The authors introduce a design-time diagnostic based on the cheapest non-identity relabelling to detect and prevent degenerate stimulus configurations before data collection In colour-based experiments, a symmetric design remained stuck even with 64x the restart budget, while an asymmetric set at the same sample size recovered
Analysis
TL;DR
- Unsupervised representational alignment is fundamentally bounded by the automorphism group of stimulus geometry, meaning more data alone cannot resolve certain symmetries
- The authors introduce a design-time diagnostic based on the cheapest non-identity relabelling to detect and prevent degenerate stimulus configurations before data collection
- In colour-based experiments, a symmetric design remained stuck even with 64x the restart budget, while an asymmetric set at the same sample size recovered perfectly every time
- Discriminating representational models and recovering correspondence are essentially uncorrelated objectives (r = -0.02 over 3,000 subsets)
- A simple diagnostic-driven colour selection of just 9 stimuli reduced catastrophic alignment failures from 75% to 2% across 93 model representations with all other factors held fixed
Why It Matters
This work reveals a fundamental identifiability limitation in unsupervised representational alignment that practitioners may unknowingly encounter when designing stimulus sets for neural representation analysis. The finding that symmetric designs (evenly spaced orientations, tones, or motion directions) create irrecoverable degeneracies challenges the common assumption that collecting more data or increasing computational budget will resolve alignment ambiguities. The proposed diagnostic offers a cheap, pre-collection safeguard that can prevent wasted experimental effort.
Technical Details
- The paper formalizes how the automorphism group of stimulus geometry constrains identifiability in unsupervised representational alignment, establishing that degeneracies exist before any data is collected
- The authors propose a design-time diagnostic using the cheapest non-identity relabelling as a measure of symmetry-induced degeneracy, transforming a known invariance (Demetci et al., 2024) into a practical intervention tool
- Experiments in colour space demonstrate the structural failure: symmetric designs with dense sampling produce near-duplicates whose transposition is nearly cost-free, rendering alignment degenerate regardless of restart budget
- The diagnostic was applied to select 9 colours without consulting any learned representation, successfully moving all 93 tested model representations away from degenerate points while holding models, layers, N, and solver constant
- The same symmetry risk generalizes to any regularly spaced stimulus dimensions including orientations, tones, or motion directions, with the diagnostic requiring only one function call before data collection
Industry Insight
- Researchers designing fMRI or neural recording experiments should screen stimulus geometries for symmetry-induced degeneracies before data collection, as the cost of a single diagnostic call is negligible compared to the expense of failed experiments
- The near-zero correlation between model discrimination and correspondence recovery suggests that benchmarking alignment methods on symmetric designs may produce misleadingly optimistic or pessimistic results depending on the specific symmetry present
- Tool developers should integrate symmetry diagnostics into stimulus design pipelines for representational similarity analysis, particularly for common regular designs like evenly spaced angular or chromatic stimuli
Disclaimer: The above content is generated by AI and is for reference only.