AIToday
Large Language ModelsAI Safety & AlignmentAI Business & IndustryTHE DECODERPublished: Aug 30, 2026, 22:01 JST2 min read

AI helps grades, not learning: study

AI helps grades, not learning: study

Key takeaway

  • GPT-4o significantly improved student grades at Bocconi University. A causal reasoning lesson increased idea diversity but not scores.

  • Whether AI improves learning remains unclear.

  • The study involved 1,053 freshmen in November 2025.

3 Key Points

  1. What happened

    A randomized experiment with 1,053 freshmen at Bocconi University found that GPT-4o helped students earn significantly better grades on a business assignment, scoring nearly a full point higher on a 1-to-5 scale.

  2. Why it matters

    The study suggests that current grading systems reward polish and structure, not understanding. Since AI can produce expert-like work, grades alone may not reflect what students actually learn.

  3. What to watch

    A short lesson on causal reasoning didn't raise traditional scores but encouraged more diverse ideas and deeper thinking. Combining it with GPT-4o didn't improve scores further, and no follow-up test measured retained knowledge.

Ask the AI about this article →

Context & Analysis

This study adds to a growing body of evidence that AI can improve task performance without improving underlying skills. A prior analysis of over 500,000 US college grades found that top grades increased most in writing and programming courses with heavy homework after ChatGPT's launch. Controlled experiments showed that when AI was removed, participants performed worse than a control group, especially those who used GPT for direct answers. A 30-month study of 26,000 students in China found homework improved but exam scores dropped, with entrance exam results 18 to 24 percent lower over the long term.

The Bocconi experiment highlights a tension in education: current grading rewards conventional, polished answers, not originality or deep reasoning. The authors argue for changing grading criteria to explicitly reward originality and reasoning. However, as the article notes, this would likely require systemic changes. The study also has limitations: it involved only freshmen at one university, randomization was at the class-section level, and some measures were evaluated by AI models, including from OpenAI, which was involved in the research. Whether the GPT advantage reflects actual learning or just better output remains an open question.

FAQ

Did the causal reasoning lesson improve grades?
No, it didn't raise traditional scores. In fact, work from those students scored slightly worse on average, but they produced more diverse ideas and explained their reasoning more.
What did the study measure besides grades?
It also measured argumentation quality, number of ideas, idea diversity, and text properties. The grading rubric correlated with higher scores for more ideas and coherence but penalized falsifiability and divergence from peers.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • AI agents have no sense of time, study findsTHE DECODER · 2h ago
  • OpenAI and METR release final reports on Hugging Face breachITmedia AI+ · 5h ago
  • Anthropic cuts Claude Code weekly limits by 17%THE DECODER · 5h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleS. Korea's aging population may blunt AI boom wealth