AI News AI资讯 21h ago Updated 18h ago 更新于 18小时前 46

The skills that earn top grades are the ones AI can fake best 获得高分的技能正是AI最能伪造的技能

GPT-4o significantly boosted grades (nearly a full point on a 1-5 scale) on a business marketing assignment among 1,053 freshmen at Bocconi University, primarily by improving content quality and logical coherence rather than student knowledge A short causal reasoning lesson did not raise traditional scores but encouraged more diverse, unconventional solutions and deeper thinking about causal mechanisms and falsifiability The two interventions (AI access and causal reasoning training) had complem 博科尼大学1053名新生的随机对照实验显示,GPT-4o帮助学生在商业作业中获得显著更高的分数(1-5分制下提升近1分),答案更连贯且更贴近专家建议 简短的因果推理课程未提高传统评分,但促使学生产生更多样化、更深入思考因果关系和假设条件的解决方案 当前评分标准奖励结构化、完整的答案,却惩罚原创性和多样性,使AI极易被用作"合规作弊"工具 研究未设置无AI辅助的后续测试,无法确定成绩提升源于学生实际知识增长还是仅AI输出质量改善 OpenAI参与研究并提供技术,作者指出教育评估需转向奖励原创性、推理过程和多元方法,而非仅看 polished 的成品

68
Hot 热度
65
Quality 质量
62
Impact 影响力

Analysis 深度分析

TL;DR

  • GPT-4o significantly boosted grades (nearly a full point on a 1-5 scale) on a business marketing assignment among 1,053 freshmen at Bocconi University, primarily by improving content quality and logical coherence rather than student knowledge
  • A short causal reasoning lesson did not raise traditional scores but encouraged more diverse, unconventional solutions and deeper thinking about causal mechanisms and falsifiability
  • The two interventions (AI access and causal reasoning training) had complementary effects: AI improved polish and conventional quality, while the lesson improved originality and reasoning depth
  • Current grading rubrics penalize divergence and originality while rewarding conventional, well-structured answers, creating a structural incentive for AI-assisted cheating rather than genuine learning
  • The study does not demonstrate improved learning—no follow-up test without AI was conducted, leaving open whether students actually retained knowledge or simply produced better outputs

Why It Matters

This research directly challenges the assumption that higher grades equal better learning in the age of AI, revealing a fundamental misalignment between current assessment systems and genuine educational outcomes. For AI practitioners and educators, it underscores that grading reform is not optional but essential—if rubrics continue to reward polish and conventionality, AI will systematically game the system while students' actual understanding erodes.

Technical Details

  • Study design: Randomized controlled experiment across 13 sections of an introductory management course at Bocconi University (November 2025), with 1,053 freshmen assigned to four groups: control, causal reasoning lesson, GPT-4o access, or both
  • Assignment: Students wrote marketing recommendations for the university's merchandise shop in up to 180 words; graded on a 1-5 scale by blind human graders, with additional metrics (causal reasoning, idea diversity) evaluated using OpenAI and Anthropic models
  • Causal reasoning intervention: A short lesson covering coherent causal logic, falsifiability, and the link between proposed actions and desired outcomes; did not improve traditional scores but increased explanation depth and idea divergence
  • GPT-4o impact: Students with AI access produced answers with ~2 more ideas on average, greater logical coherence, and closer alignment with subject-matter expert recommendations; the advantage persisted even after controlling for argumentation quality, idea count, diversity, and text properties
  • Grading rubric findings: Stronger falsifiability, detailed causal explanations, and divergence from peers' ideas correlated with lower scores, while idea quantity, coherence, and conventional alignment correlated with higher scores

Industry Insight

  • Assessment redesign is urgent: Institutions must explicitly reward originality, reasoning depth, and consideration of multiple approaches in grading criteria—otherwise AI will continue to inflate grades without corresponding learning gains
  • AI as crutch vs. scaffold matters: Evidence from multiple studies (including a 30-month Chinese study showing 18-24% exam score drops) suggests the critical distinction is whether AI replaces student thinking or supports it; educators should design assignments that require AI-verified reasoning rather than AI-generated answers
  • OpenAI's dual role raises concerns: With OpenAI providing the technology and several authors affiliated with the company, independent replication and critical scrutiny of the framing (which emphasizes grading reform over learning outcomes) are warranted before widespread policy adoption

TL;DR

  • 博科尼大学1053名新生的随机对照实验显示,GPT-4o帮助学生在商业作业中获得显著更高的分数(1-5分制下提升近1分),答案更连贯且更贴近专家建议
  • 简短的因果推理课程未提高传统评分,但促使学生产生更多样化、更深入思考因果关系和假设条件的解决方案
  • 当前评分标准奖励结构化、完整的答案,却惩罚原创性和多样性,使AI极易被用作"合规作弊"工具
  • 研究未设置无AI辅助的后续测试,无法确定成绩提升源于学生实际知识增长还是仅AI输出质量改善
  • OpenAI参与研究并提供技术,作者指出教育评估需转向奖励原创性、推理过程和多元方法,而非仅看 polished 的成品

为什么值得看

这项研究揭示了AI辅助学习与传统教育评估体系之间的根本矛盾:当AI能轻松生成高分作业时,现有评分机制实际上在奖励"AI能模仿的技能"而非"学生真正掌握的能力"。对教育科技从业者、政策制定者和高校管理者而言,这不仅是关于如何整合AI工具的教学问题,更是关于如何重新定义学习成果评估的系统性挑战。

技术解析

  • 实验设计:2025年11月,博科尼大学13个管理入门课程班级被随机分配到四组——对照组、因果推理课程组、GPT-4o访问组、以及两者结合组。学生需为大学商品店撰写最多180字的营销建议。因果推理课程涵盖连贯的因果逻辑、可证伪性,以及提议行动如何导向预期结果。
  • 评估方法:主要成绩由不知情的真人评分员评定;因果推理和想法多样性等辅助指标由OpenAI和Anthropic的模型评估。即使控制论证质量、想法数量、多样性及文本特征后,GPT-4o的优势依然显著。
  • GPT-4o效果:实验组平均得分提升近1分(1-5分制),答案包含约多两个想法、逻辑更连贯,且更贴近三位学科专家的建议。研究者认为这反映的是内容质量提升而非学生知识增长。
  • 因果推理课程效果:传统分数略有下降,但学生更常解释建议为何有效、在何种条件下可能失败,并产生更多偏离同伴的多样化想法。与GPT-4o结合未进一步提升传统分数,但在因果推理指标上形成互补。
  • 研究局限:无AI辅助的后续测试验证学习保留;仅针对单一大学新生和狭窄营销任务;随机化在班级层面而非个人层面;OpenAI提供技术并参与研究,存在利益关联。

行业启示

  • 教育评估体系亟需结构性改革:当前以"成品质量"为核心的评分标准天然有利于AI辅助,必须将原创性、推理过程和多元视角明确纳入评估维度,否则AI将加剧教育不平等并削弱学习成效。高校和考试机构应重新设计作业与考核方式,例如引入口头答辩、过程性评估或限时无AI考试。
  • AI整合策略需区分"替代性使用"与"支持性使用":研究与其他大规模数据(美国50万学生成绩、中国2.6万学生30个月追踪)一致表明,用AI直接获取答案会导致长期表现下降,而将AI作为思维脚手架则能维持学习效果。教育机构应明确界定并引导学生采用支持性使用模式,而非简单禁止或放任。
  • AI教育产品设计应从"生成完美答案"转向"促进深度思考":OpenAI等厂商已意识到评估危机并提出改革方向,但系统性变革涉及整个教育体系。AI工具的设计应集成因果推理训练、多样性激发、可证伪性检验等功能,帮助学生在AI辅助下真正提升认知能力,而非仅优化作业输出。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

GPT GPT Education AI 教育AI Research 科学研究 Evaluation 评测