GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus
GLAN-QnA-KR is a 303,581-row synthetic Korean instruction-QA corpus generated using Microsoft's Phi-3.5-MoE-instruct model via a seedless taxonomy-driven pipeline. The dataset covers 1,084 disciplines with English labels and Korean text, featuring a difficulty scale of 100-900 and high uniqueness with near-zero duplicate clusters. Rigorous contamination audits against major benchmarks (KMMLU, KoBEST, HAE-RAE-Bench) show minimal overlap, ensuring data integrity for training. It stands as the larg
Analysis
TL;DR
- GLAN-QnA-KR is a 303,581-row synthetic Korean instruction-QA corpus generated using Microsoft's Phi-3.5-MoE-instruct model via a seedless taxonomy-driven pipeline.
- The dataset covers 1,084 disciplines with English labels and Korean text, featuring a difficulty scale of 100-900 and high uniqueness with near-zero duplicate clusters.
- Rigorous contamination audits against major benchmarks (KMMLU, KoBEST, HAE-RAE-Bench) show minimal overlap, ensuring data integrity for training.
- It stands as the largest single-pipeline synthetic Korean instruction corpus verifiable on Hugging Face Hub under a seedless protocol.
- Released under OpenRAIL license, it provides an openly redistributable resource for advancing Korean NLP capabilities.
Why It Matters
This corpus addresses a critical gap in high-quality, large-scale synthetic data for Korean language models, which often lag behind English resources. By providing a rigorously audited, contamination-free dataset, it enables researchers to train more robust and accurate Korean LLMs without the risk of data leakage from evaluation sets. This facilitates better performance in downstream tasks such as instruction tuning and domain-specific knowledge acquisition.
Technical Details
- Generation Pipeline: Utilizes the GLAN synthesis pipeline with Microsoft's Phi-3.5-MoE-instruct as the producer model, generating data in December 2024.
- Dataset Structure: Contains 303,581 rows spanning 1,084 English-labeled disciplines, with Korean Q&A pairs, a median question length of 313 characters, and answer length of 1,098 characters.
- Uniqueness Metrics: Exact duplicates are extremely rare (1 in 303,581), and character-trigram near-duplicate clusters (Jaccard >= 0.9) are zero in a 5,000-sample probe.
- Contamination Audit: Tested against KMMLU, KoBEST, and HAE-RAE-Bench; maximum character-trigram Jaccard similarity was 0.163, with no test items exceeding 0.7. Multilingual-E5 cosine similarity showed only one item at >= 0.90 and none at >= 0.95.
- Licensing: Distributed under the OpenRAIL license, allowing for open redistribution and research use.
Industry Insight
- Data Quality over Quantity: The emphasis on low contamination and high uniqueness highlights the importance of rigorous auditing in synthetic data generation, setting a new standard for Korean NLP datasets.
- Scalability of Seedless Methods: The success of the seedless taxonomy-driven approach suggests that automated, scalable methods can produce high-quality instructional data without manual seed curation, reducing costs and increasing coverage.
- Resource Availability: As the largest verified synthetic Korean corpus, GLAN-QnA-KR will likely become a foundational resource for pre-training and fine-tuning Korean LLMs, accelerating progress in low-resource language AI.
Disclaimer: The above content is generated by AI and is for reference only.