Ask HN: Can you still tell AI-generated text apart in your own language?
AI-generated Japanese text can still be detected through unnatural word choices, with formal terms like 実務 and 帳簿 appearing far too frequently, while casual expressions like 効く and 刺さる are misapplied AI-generated English has improved significantly and is becoming harder for native speakers to distinguish from human writing Model selection matters substantially for non-English languages, with Anthropic's Opus 4.6 and Google's models outperforming OpenAI's offerings in Japanese long-form writing G
Analysis
TL;DR
- AI-generated Japanese text can still be detected through unnatural word choices, with formal terms like 実務 and 帳簿 appearing far too frequently, while casual expressions like 効く and 刺さる are misapplied
- AI-generated English has improved significantly and is becoming harder for native speakers to distinguish from human writing
- Model selection matters substantially for non-English languages, with Anthropic's Opus 4.6 and Google's models outperforming OpenAI's offerings in Japanese long-form writing
- Gmail's AI drafting feature demonstrated deep cultural absorption of Japanese corporate email conventions, marking a notable milestone in multilingual AI capabilities
Why It Matters
This observation highlights a critical gap in multilingual AI evaluation — while English-language AI quality has reached near-human levels, non-English outputs still carry detectable artifacts that reveal their origin. For AI practitioners and researchers, this underscores the need for language-specific quality benchmarks and the importance of evaluating models in users' native languages rather than assuming English performance generalizes.
Technical Details
- Japanese AI text exhibits distinctive lexical biases: overuse of formal/bureaucratic vocabulary (実務, 帳簿) and misapplication of casual verbs (効く, 刺さる), analogous to the English "delve" phenomenon
- Gmail's AI drafting feature demonstrated contextual cultural absorption, internalizing Japanese corporate email norms and bureaucratic register conventions
- Model-level performance varies significantly by language: Anthropic Opus 4.6 and Google models produce competent long-form Japanese, while OpenAI models (Opus 4.8/5, Fable) produce text described as overly "geeky" in tone
- Chinese, expected to benefit from heavy training data, remains an open question alongside Spanish, Hindi, and other widely spoken languages
Industry Insight
- AI companies should invest in native-speaker evaluation for non-English languages rather than relying on English-centric benchmarks, as subtle linguistic artifacts persist even as gross errors disappear
- The "model matters more than language" insight suggests that multilingual capability is not uniform across providers, and organizations should benchmark models per-language for their specific use cases
- Cultural register absorption (as seen in Gmail's Japanese email drafting) represents a frontier beyond translation accuracy — future competitive advantage will come from models that internalize domain-specific and culturally nuanced communication conventions
Disclaimer: The above content is generated by AI and is for reference only.