Research Papers 论文研究 3h ago Updated 1h ago 更新于 1小时前 39

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States 美国网络流媒体录音宗教广播转录本的大型语料库

A large-scale corpus of transcribed religious radio broadcasts from US webstreams was created, covering over 700,000 recordings and more than 60 million diarized speech lines. The data was collected from 785 distinct streams rebroadcasting signals from over two thousand AM and FM stations during July 2025. Automated transcription and speaker diarization were applied, with segmentation and labeling by format and topic using a large language model. The dataset supports research in religious broadc 该研究构建了一个大规模的美国宗教广播转录语料库,填补了该领域缺乏大型文本数据的空白。 通过自动化管道对785个网络流媒体进行录音、转录和说话人分离,生成了超过6000万行带时间戳的语音文本。 利用大语言模型对录音内容按节目格式和主题进行了自动标注与分段。 该数据集支持跨地区、跨传统的宗教传播描述性研究,以及社会政治议题在宗教媒体中的话语分析。 为语音处理领域提供了一个具有代表性的低资源(underrepresented)数据子集,有助于推动相关模型的泛化能力。

55
Hot 热度
60
Quality 质量
55
Impact 影响力

Analysis 深度分析

TL;DR

  • A large-scale corpus of transcribed religious radio broadcasts from US webstreams was created, covering over 700,000 recordings and more than 60 million diarized speech lines.
  • The data was collected from 785 distinct streams rebroadcasting signals from over two thousand AM and FM stations during July 2025.
  • Automated transcription and speaker diarization were applied, with segmentation and labeling by format and topic using a large language model.
  • The dataset supports research in religious broadcasting analysis, social/political discourse in religious media, and speech processing in underrepresented domains.

Why It Matters

This dataset fills a critical gap in available resources for studying religious media content at scale, enabling new forms of computational analysis that were previously constrained by lack of transcript data. For NLP and speech researchers, it provides a rich domain-specific corpus to develop and evaluate models tailored to religious broadcast characteristics, potentially improving performance in similar understudied areas. Sociologists and communication scholars gain unprecedented access to longitudinal, geographically diverse religious programming content for analyzing trends in messaging across traditions and regions.

Technical Details

  • Data collection involved automated capture of 15-minute segments from 785 live webstreams representing over 2,000 traditional radio stations across the United States during July 2025
  • Processing pipeline included automatic speech recognition for transcription followed by speaker diarization to distinguish different speakers within each recording
  • Large language model was used to segment recordings into meaningful units and label them according to programming format (e.g., sermon, talk show, music) and topical categories
  • Final output structured as three linked tables: stream metadata (source information), recording metadata (timestamps, duration, quality metrics), and transcript lines (text content with speaker identities and timing annotations)
  • Total volume comprises approximately 700,000 individual recordings containing more than 60 million lines of diarized speech text

Industry Insight

The availability of this specialized religious broadcast corpus creates opportunities for developing domain-adapted language models that understand theological terminology, rhetorical patterns specific to preaching styles, and contextual nuances in faith-based discussions. Media monitoring companies could leverage such datasets to track evolving perspectives on social issues within religious communities, providing valuable insights for political strategists or nonprofit organizations targeting faith audiences. Additionally, speech technology developers might use this resource to improve accessibility tools like real-time captioning services for religious programming, enhancing inclusivity for hearing-impaired congregants who consume content through digital platforms rather than traditional radio receivers.

TL;DR

  • 该研究构建了一个大规模的美国宗教广播转录语料库,填补了该领域缺乏大型文本数据的空白。
  • 通过自动化管道对785个网络流媒体进行录音、转录和说话人分离,生成了超过6000万行带时间戳的语音文本。
  • 利用大语言模型对录音内容按节目格式和主题进行了自动标注与分段。
  • 该数据集支持跨地区、跨传统的宗教传播描述性研究,以及社会政治议题在宗教媒体中的话语分析。
  • 为语音处理领域提供了一个具有代表性的低资源(underrepresented)数据子集,有助于推动相关模型的泛化能力。

为什么值得看

对于从事自然语言处理、社会计算或媒体研究的从业者而言,此数据集提供了罕见且结构化的真实世界宗教话语样本,可用于训练或评估面向特定文化语境的语言模型。同时,其多模态元数据结构也为构建可解释的媒体分析系统提供了基础支撑。

技术解析

数据采集覆盖2025年7月期间美国境内超过2000个AM/FM电台的785个独立网络流,每段15分钟滚动录制,累计生成超70万次音频片段。所有音频均经过端到端自动流水线处理:首先使用ASR系统进行语音识别并输出带说话人分离(diarization)的文本序列,随后由大语言模型依据上下文对每条记录进行分类标签分配(如福音派、灵恩派、教义讲解等),最终输出结构化表格形式的数据集,包含流源信息、录音元数据及逐行转录文本三类关联表。

行业启示

宗教媒体作为传统大众传播的重要分支,在数字时代仍保有广泛受众,但其数字化表征严重不足;本工作展示了如何利用现代NLP技术从非结构化广播内容中提取高价值语义信息,为后续建立动态舆情监测框架提供范本。此外,此类垂直领域语料的积累将促进AI系统在多元文化场景下的公平性与包容性设计,尤其在涉及信仰、价值观敏感话题时避免偏见放大。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Dataset 数据集 Research 科学研究