On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?
The study evaluates four proprietary LLMs (GPT-4o, GPT-5.2, Claude Sonnet 4.5, DeepSeek) for specialized terminology translation from English to French in two domains: Earth, Environmental and Planetary Sciences (EEPS) and Natural Language Processing (NLP). Two prompting strategies were tested: terminology mode and translation mode, with 80 terms per domain. Claude Sonnet 4.5 achieved the best results in the most favorable configuration, while DeepSeek demonstrated greater stability across tests
Analysis
TL;DR
- The study evaluates four proprietary LLMs (GPT-4o, GPT-5.2, Claude Sonnet 4.5, DeepSeek) for specialized terminology translation from English to French in two domains: Earth, Environmental and Planetary Sciences (EEPS) and Natural Language Processing (NLP).
- Two prompting strategies were tested: terminology mode and translation mode, with 80 terms per domain.
- Claude Sonnet 4.5 achieved the best results in the most favorable configuration, while DeepSeek demonstrated greater stability across tests.
- Confidence estimates provided by models were only a partial indicator of terminological accuracy.
- LLMs can assist specialized translators but cannot currently replace specialized corpora.
Why It Matters
This research is highly relevant to AI practitioners and researchers working on natural language processing, particularly those focused on machine translation and terminology management. It provides empirical evidence on the capabilities and limitations of current LLMs in specialized domains, which can inform both tool development and practical applications in professional translation workflows. The findings also highlight the importance of continued improvement in model reliability and confidence estimation mechanisms.
Technical Details
- Four proprietary LLMs were evaluated: GPT-4o, GPT-5.2, Claude Sonnet 4.5, and DeepSeek.
- Testing was conducted in two specialized domains: Earth, Environmental and Planetary Sciences (EEPS) and Natural Language Processing (NLP).
- Each domain included 80 test terms, totaling 160 terms across both fields.
- Two prompting strategies were compared: terminology-focused prompts and translation-context prompts.
- Performance metrics likely included accuracy of term equivalents, consistency across domains, and correlation between model confidence scores and actual correctness.
- The study emphasizes that while some models perform well under optimal conditions, none consistently match the reliability of curated specialized corpora.
Industry Insight
Specialized translation professionals should consider integrating LLMs as supplementary tools rather than replacements for established terminological resources like domain-specific corpora or glossaries. Model selection matters—Claude Sonnet 4.5 shows promise in high-performance scenarios, while DeepSeek offers more predictable behavior, making it suitable for consistent output needs. Future developments should focus on improving confidence calibration and reducing variability across domains to make LLMs more trustworthy for critical translation tasks.
Disclaimer: The above content is generated by AI and is for reference only.