ChatGPT Raised Assignment Scores. The Bocconi Study Did Not Measure Learning

A 1,053-student trial found that ChatGPT improved scored recommendations, but its originality metric was computational and the study did not measure later learning.

Share this article

Giving students ChatGPT access raised the quality score on a tightly specified business assignment. A short causal-reasoning exercise did something different: it pushed students toward more varied ideas that the assignment’s standard rubric did not reward.

That is the useful result of a new experiment involving 1,053 first-year students. It is not evidence that ChatGPT improved learning. The trial measured one assisted output produced in about 45 minutes; it did not test what students understood, retained, or could transfer to a later task without the model.

The practical lesson is about measurement. An assessment that records only the finished answer can reward conventional polish while missing originality, reasoning, and independent understanding. Educators—and teams evaluating AI-assisted knowledge work—need separate measures for those outcomes.

What the ChatGPT critical-thinking study randomized

The August working paper describes a preregistered 2×2 randomized controlled trial conducted in November 2025. Researchers assigned 13 intact class sections within three degree programs to one of four conditions. This was cluster randomization at the class level, not individual random assignment.

ArmNInterventionResult
Control249Placebo game; no ChatGPTBaseline
Causal training256Reasoning game; no ChatGPTMore diverse ideas; no higher rubric score
ChatGPT197GPT-4o availableEstimated rubric score rose 0.862 points
Both351Training and GPT-4o availableRetained much of each distinct pattern

All students then received the same business case: write roughly 180 words recommending ways to increase alumni awareness and use of the university’s merchandise shop. Twenty trained master’s students graded each response, with three raters per submission, on a five-point rubric derived from two marketing concepts—awareness and usage. Three domain experts also supplied reference recommendations.

The assignment unit matters. Only 13 classes—not 1,053 individual students—were randomized. The authors clustered their standard errors and added controls for variables that failed balance checks, but the design still has far fewer independent assignment units than the student count in the headline suggests. Among participants given GPT access, 90.5% reported complying with their assigned treatment.

The OpenAI account and Bocconi’s report identify the university as Bocconi. The paper instead calls it a leading European university and withholds both its name and the American Economic Association trial-registry number to preserve anonymity. The appendix later includes Bocconi-specific task text. This is a documentation and reproducibility gap, not evidence of a different trial site.

The 0.862-point gain belongs to one rubric

ChatGPT access increased the estimated awareness-and-usage score by 0.862 points relative to an estimated control score of 2.09. On a five-point scale, that is a substantial effect for the tested task. The ChatGPT-assisted recommendations also moved closer to the three expert texts.

The paper’s statistical analysis attributes part of the gain to more ideas, clearer logic, and other textual features. Roughly a third of the effect remained after its fullest set of text controls. The researchers interpret that remainder as an improvement in the substance of the recommendations, not merely their presentation.

That finding still has a hard boundary: “substance” here means a stronger recommendation under the awareness-and-usage evaluation used in the experiment. The trial did not test factual recall, durable business knowledge, independent reasoning after ChatGPT was removed, or performance on a new kind of problem. A higher assignment score cannot carry those claims.

The intervention also used GPT-4o through ChatGPT Edu. Results from one model, account environment, course, and assignment should not be transferred to a newer model or another discipline without another test.

Originality was a computed distance, not a creativity grade

The causal-reasoning exercise taught students to build cause-and-effect chains, state conditions that could falsify a proposal, and identify mechanisms that would make it work. Students who received it used more mechanism-based and falsifiable reasoning. Their ideas were also more diverse within each answer and more different from those of other students.

The standard marketing rubric did not reward those changes. In the paper’s models, answers farther from the typical solution space could receive lower scores. This explains why the training could change the reasoning and the idea set without improving the headline assignment grade.

“Originality” needs its own caveat. Human graders did not directly score creative merit. The appendix describes a multi-stage computational pipeline:

  1. GPT-5.2 extracted candidate ideas repeatedly under three prompt variants.
  2. Repeated runs were reconciled, with uncertain matches judged by GPT-5.2 and Claude Opus 4.6.
  3. Confirmed ideas were embedded with OpenAI’s text-embedding-3-large model.
  4. Cosine distance measured how far ideas were from one another within and across submissions.

This is a disclosed and fairly elaborate operational measure of semantic diversity. It can detect difference in the text; it does not establish that an idea is useful, feasible, surprising to a domain expert, or genuinely new outside this class. Because OpenAI models participate in the extraction, judging, and embedding pipeline, an independent replication should also test whether the result survives different models and human originality ratings.

Did the experiment show that students learned more?

No. It showed that ChatGPT access improved a finished recommendation under a specified rubric, while causal-reasoning training improved other measured properties of that recommendation.

Learning requires another observation after the assistance ends. For example, an evaluator could ask students to explain their causal model, critique an unfamiliar proposal, reproduce the reasoning after a delay, or solve a structurally different problem without ChatGPT. The Bocconi experiment did none of those things.

That missing stage matters because assisted performance and learning can move in different directions. In a separate peer-reviewed field experiment involving high-school mathematics, researchers found that GPT-4 access improved practice performance, but students using an unrestricted assistant later scored 17% lower than the control group on an unassisted exam. A tutor designed to provide hints rather than answers largely mitigated that loss. The PNAS paper tested a different age group, subject, tool design, and outcome, so it is not a replication or rebuttal. It demonstrates why an assisted artifact cannot substitute for a later learning measure.

The new study is currently a working paper distributed through OpenAI’s channels. The source scan on August 28 found no peer-reviewed publication or independent replication. A separate reconstruction by Next AI Press reached the same main scope warning: the result concerns a single task and computational measures, not learning across a course.

OpenAI also participated in the research. The author list includes current OpenAI researchers, a contributor who worked on the project while employed there, and a paid OpenAI contractor. Those affiliations do not invalidate the experiment, but they make independent replication and a public preregistration record especially valuable.

The study’s most important result is the missing learning measure

The trial separated output quality, causal reasoning, and a computational originality measure well enough to show that the treatments moved them differently. It did not observe what students could retain or transfer after ChatGPT was removed.

That omission does not invalidate the reported 0.862-point output gain. It limits the claim to an assisted artifact produced under one rubric. A delayed, unfamiliar, unassisted task would answer a different and more educationally consequential question.

The study therefore supports a modest conclusion: ChatGPT helped students produce stronger recommendations in this setting, while the causal-reasoning intervention changed aspects of the work the main rubric did not reward. It does not show that either group became better independent thinkers.

Sources

  1. Training novices to think, or giving them LLMs? Evidence from an RCT
  2. OpenAI: Better answers, broader thinking
  3. Bocconi University: Better evaluations with ChatGPT
  4. Next AI Press reconstruction of the ChatGPT student study
  5. PNAS: Generative AI without guardrails can harm learning