The skills that earn top grades are the ones AI can fake best
GPT-4o significantly boosted grades (nearly a full point on a 1-5 scale) on a business marketing assignment among 1,053 freshmen at Bocconi University, primarily by improving content quality and logical coherence rather than student knowledge A short causal reasoning lesson did not raise traditional scores but encouraged more diverse, unconventional solutions and deeper thinking about causal mechanisms and falsifiability The two interventions (AI access and causal reasoning training) had complem
Analysis
TL;DR
- GPT-4o significantly boosted grades (nearly a full point on a 1-5 scale) on a business marketing assignment among 1,053 freshmen at Bocconi University, primarily by improving content quality and logical coherence rather than student knowledge
- A short causal reasoning lesson did not raise traditional scores but encouraged more diverse, unconventional solutions and deeper thinking about causal mechanisms and falsifiability
- The two interventions (AI access and causal reasoning training) had complementary effects: AI improved polish and conventional quality, while the lesson improved originality and reasoning depth
- Current grading rubrics penalize divergence and originality while rewarding conventional, well-structured answers, creating a structural incentive for AI-assisted cheating rather than genuine learning
- The study does not demonstrate improved learning—no follow-up test without AI was conducted, leaving open whether students actually retained knowledge or simply produced better outputs
Why It Matters
This research directly challenges the assumption that higher grades equal better learning in the age of AI, revealing a fundamental misalignment between current assessment systems and genuine educational outcomes. For AI practitioners and educators, it underscores that grading reform is not optional but essential—if rubrics continue to reward polish and conventionality, AI will systematically game the system while students' actual understanding erodes.
Technical Details
- Study design: Randomized controlled experiment across 13 sections of an introductory management course at Bocconi University (November 2025), with 1,053 freshmen assigned to four groups: control, causal reasoning lesson, GPT-4o access, or both
- Assignment: Students wrote marketing recommendations for the university's merchandise shop in up to 180 words; graded on a 1-5 scale by blind human graders, with additional metrics (causal reasoning, idea diversity) evaluated using OpenAI and Anthropic models
- Causal reasoning intervention: A short lesson covering coherent causal logic, falsifiability, and the link between proposed actions and desired outcomes; did not improve traditional scores but increased explanation depth and idea divergence
- GPT-4o impact: Students with AI access produced answers with ~2 more ideas on average, greater logical coherence, and closer alignment with subject-matter expert recommendations; the advantage persisted even after controlling for argumentation quality, idea count, diversity, and text properties
- Grading rubric findings: Stronger falsifiability, detailed causal explanations, and divergence from peers' ideas correlated with lower scores, while idea quantity, coherence, and conventional alignment correlated with higher scores
Industry Insight
- Assessment redesign is urgent: Institutions must explicitly reward originality, reasoning depth, and consideration of multiple approaches in grading criteria—otherwise AI will continue to inflate grades without corresponding learning gains
- AI as crutch vs. scaffold matters: Evidence from multiple studies (including a 30-month Chinese study showing 18-24% exam score drops) suggests the critical distinction is whether AI replaces student thinking or supports it; educators should design assignments that require AI-verified reasoning rather than AI-generated answers
- OpenAI's dual role raises concerns: With OpenAI providing the technology and several authors affiliated with the company, independent replication and critical scrutiny of the framing (which emphasizes grading reform over learning outcomes) are warranted before widespread policy adoption
Disclaimer: The above content is generated by AI and is for reference only.