Mimicry without understanding: the origins of decision bias in large language models
LLMs can generate biased decisions through two mechanisms even when training data lacks bias: faulty mimicry (inferring preferences from unrelated behaviors) and mimicry of explicitly biased descriptions ChatGPT-4o and Qwen exhibited social proof biases when prompted with human behaviors that were clearly non-indicative of actual preferences LLMs displayed loss aversion when it was explicitly described as a bias in scientific reports, with the degree of reported bias predicting the model's own b
Analysis
TL;DR
- LLMs can generate biased decisions through two mechanisms even when training data lacks bias: faulty mimicry (inferring preferences from unrelated behaviors) and mimicry of explicitly biased descriptions
- ChatGPT-4o and Qwen exhibited social proof biases when prompted with human behaviors that were clearly non-indicative of actual preferences
- LLMs displayed loss aversion when it was explicitly described as a bias in scientific reports, with the degree of reported bias predicting the model's own biased responses
- Scientific papers documenting biases can become self-fulfilling prophecies in LLM outputs
- The study reveals underlying component processes driving bias generation beyond simply cataloging existing LLM biases
Why It Matters
This research is critically relevant to AI practitioners because it demonstrates that bias in LLMs is not solely a product of biased training data—models can generate biases through flawed inference and mimicry of biased descriptions alone. For researchers, it opens new avenues for understanding the cognitive mechanisms behind LLM decision-making and highlights the risk that academic literature on biases may inadvertently propagate those same biases into model outputs.
Technical Details
- The study examined two bias generation mechanisms: (1) faulty mimicry of preferences based on human behavior where LLMs infer preferences from logically unrelated behaviors, and (2) mimicry of explicitly biased human behaviors
- Four studies were conducted focusing on economic biases, testing ChatGPT-4o and Qwen models
- Social proof bias was observed even when prompted with reports of human behaviors clearly non-indicative of individuals' actual preferences
- Loss aversion bias emerged when explicitly described in scientific reports, with the extent of bias in the report predicting the LLM's subsequent biased responses
- The research falls under computation and language (cs.CL), artificial intelligence (cs.AI), and human-computer interaction (cs.HC)
Industry Insight
- AI developers should implement bias detection that goes beyond training data auditing, as models can generate novel biases through inference mechanisms independent of their training distribution
- Prompt engineering guidelines should account for the risk that descriptive text about biases can activate those same biases in model outputs, creating a feedback loop
- The self-fulfilling prophecy effect suggests that bias mitigation strategies need to address not only data curation but also the models' tendency to mirror biased descriptions presented in prompts, warranting new evaluation benchmarks that test bias generation through mimicry mechanisms
Disclaimer: The above content is generated by AI and is for reference only.