A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States
A large-scale corpus of transcribed religious radio broadcasts from US webstreams was created, covering over 700,000 recordings and more than 60 million diarized speech lines. The data was collected from 785 distinct streams rebroadcasting signals from over two thousand AM and FM stations during July 2025. Automated transcription and speaker diarization were applied, with segmentation and labeling by format and topic using a large language model. The dataset supports research in religious broadc
Analysis
TL;DR
- A large-scale corpus of transcribed religious radio broadcasts from US webstreams was created, covering over 700,000 recordings and more than 60 million diarized speech lines.
- The data was collected from 785 distinct streams rebroadcasting signals from over two thousand AM and FM stations during July 2025.
- Automated transcription and speaker diarization were applied, with segmentation and labeling by format and topic using a large language model.
- The dataset supports research in religious broadcasting analysis, social/political discourse in religious media, and speech processing in underrepresented domains.
Why It Matters
This dataset fills a critical gap in available resources for studying religious media content at scale, enabling new forms of computational analysis that were previously constrained by lack of transcript data. For NLP and speech researchers, it provides a rich domain-specific corpus to develop and evaluate models tailored to religious broadcast characteristics, potentially improving performance in similar understudied areas. Sociologists and communication scholars gain unprecedented access to longitudinal, geographically diverse religious programming content for analyzing trends in messaging across traditions and regions.
Technical Details
- Data collection involved automated capture of 15-minute segments from 785 live webstreams representing over 2,000 traditional radio stations across the United States during July 2025
- Processing pipeline included automatic speech recognition for transcription followed by speaker diarization to distinguish different speakers within each recording
- Large language model was used to segment recordings into meaningful units and label them according to programming format (e.g., sermon, talk show, music) and topical categories
- Final output structured as three linked tables: stream metadata (source information), recording metadata (timestamps, duration, quality metrics), and transcript lines (text content with speaker identities and timing annotations)
- Total volume comprises approximately 700,000 individual recordings containing more than 60 million lines of diarized speech text
Industry Insight
The availability of this specialized religious broadcast corpus creates opportunities for developing domain-adapted language models that understand theological terminology, rhetorical patterns specific to preaching styles, and contextual nuances in faith-based discussions. Media monitoring companies could leverage such datasets to track evolving perspectives on social issues within religious communities, providing valuable insights for political strategists or nonprofit organizations targeting faith audiences. Additionally, speech technology developers might use this resource to improve accessibility tools like real-time captioning services for religious programming, enhancing inclusivity for hearing-impaired congregants who consume content through digital platforms rather than traditional radio receivers.
Disclaimer: The above content is generated by AI and is for reference only.