A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation
Developed a 17.62-hour speech corpus for Mizo, a low-resource language, enabling ASR system development Whisper-large-v3 achieved the best conventional WER of 18.08%, with morphology-aware evaluation dropping it to 7.22% SraVaani 1.0 zero-shot WER was 58.27%, but fine-tuning on curated Mizo data reduced it to 29.45% conventional and 17.93% morphology-aware Morphology-aware evaluation significantly outperforms conventional WER across all models, highlighting the importance of linguistic structure
Analysis
TL;DR
- Developed a 17.62-hour speech corpus for Mizo, a low-resource language, enabling ASR system development
- Whisper-large-v3 achieved the best conventional WER of 18.08%, with morphology-aware evaluation dropping it to 7.22%
- SraVaani 1.0 zero-shot WER was 58.27%, but fine-tuning on curated Mizo data reduced it to 29.45% conventional and 17.93% morphology-aware
- Morphology-aware evaluation significantly outperforms conventional WER across all models, highlighting the importance of linguistic structure in low-resource ASR
- The study demonstrates that both general multilingual models (Whisper) and Indic-specific models (SraVaani 1.0) can be effectively adapted to unseen low-resource languages with targeted data curation and fine-tuning
Why It Matters
This research addresses a critical gap in low-resource language ASR, showing that even languages with minimal digital representation can achieve competitive recognition performance through careful data collection and morphology-aware evaluation. For AI practitioners working on underrepresented languages, it provides a replicable pipeline for corpus development, model fine-tuning, and more linguistically meaningful evaluation metrics.
Technical Details
- Corpus: 17.62 hours of curated Mizo speech data collected and prepared for ASR training and evaluation
- Models evaluated: Three Whisper multilingual variants (with Whisper-large-v3 performing best) and SraVaani 1.0 Indic multilingual model
- Evaluation methodology: Dual evaluation framework using both conventional Word Error Rate (WER) and morphology-aware WER, which accounts for the agglutinative and morphologically rich nature of Mizo
- Performance results: Whisper-large-v3 — 18.08% conventional WER / 7.22% morphology-aware WER; SraVaani 1.0 zero-shot — 58.27% conventional WER; SraVaani 1.0 fine-tuned — 29.45% conventional WER / 17.93% morphology-aware WER
- Key insight: Morphology-aware evaluation consistently yields substantially lower error rates, revealing that conventional WER may overestimate the difficulty of recognizing morphologically complex low-resource languages
Industry Insight
- Morphology-aware evaluation should be adopted as a standard practice for low-resource and agglutinative languages, as conventional WER can mislead practitioners about true system performance
- Fine-tuning Indic-specific multilingual models like SraVaani 1.0 on curated regional language data offers a cost-effective path to high-quality ASR, bridging the gap between zero-shot and fully adapted systems
- The 17.62-hour corpus demonstrates that meaningful ASR systems for low-resource languages are achievable without massive data investments, encouraging similar efforts for other underrepresented languages in the Indic and global south contexts
Disclaimer: The above content is generated by AI and is for reference only.