Transparent editorial process
We combine current product documentation, public pricing, and editorial analysis. Hands-on experience is only claimed when the article identifies the workflow or evidence. Sponsorships and related products are disclosed.
Read our review methodologyRecent changes
Researchers reported a randomized experiment involving 1,053 first-year university students
Students completed the same marketing problem under four different conditions
ChatGPT access raised scores on the assignment’s standard rubric
The reported increase was about 0.86 points from an estimated control score of 2.09 on a five-point scale
Causal-reasoning training produced a wider variety of ideas
Those less-conventional ideas were not rewarded by the standard rubric
A new education study offers a more useful answer than the familiar argument over whether students should be allowed to use AI.
ChatGPT helped students produce work that looked more expert and earned higher scores. A short lesson in causal reasoning did something different: it encouraged students to develop ideas that were less like everyone else’s and to explain more clearly why those ideas might work—or fail.
When students received both, the benefits did not cancel each other out.
The practical lesson is not “AI makes students smarter,” and it is not “ban AI to protect thinking.” It is that polished output and independent reasoning are different learning outcomes. If an assignment rewards only the first, increasingly capable AI can make the grade reveal less about what a student understands.
What the researchers did
Researchers working with Bocconi University and OpenAI ran a randomized controlled experiment with 1,053 first-year students studying economics, finance and management.
Every student completed a business problem: propose ways to increase awareness and use of Bocconi-branded merchandise among alumni. Classes were assigned to one of four conditions:
- No ChatGPT and no causal-reasoning training.
- Access to ChatGPT Edu using GPT-4o.
- A short causal-reasoning exercise without ChatGPT.
- Both ChatGPT access and the reasoning exercise.
The reasoning lesson did not teach students how to prompt AI. It used a game, examples, questions and feedback to help them connect causes and effects, question assumptions and explain why a proposed action might or might not produce the intended result.
Human graders assessed the submissions on a five-point rubric focused on two conventional marketing goals. Researchers also used text analysis to examine the number and variety of ideas, signs of causal reasoning and similarity to answers written by three experts.
That design matters. It let the researchers separate help from the AI tool, help from thinking instruction and the result of combining both.
ChatGPT improved the assessed product
According to Bocconi University’s summary, ChatGPT access raised scores by about 0.86 points compared with an estimated control-group score of 2.09 on the five-point scale.
Students with AI access generated more ideas, wrote with greater coherence and produced recommendations that looked more like expert answers. The reported advantage was not explained entirely by cleaner writing or a longer list of suggestions; the analysis attributed part of it to stronger substance within the terms of the assignment.
This is meaningful. A novice given a tool can close part of the gap between their work and an expert example. For students still learning a professional format, seeing a better structure and a broader set of conventional options may provide useful scaffolding.
But a stronger submission is not the same thing as stronger independent mastery. The experiment assessed the work students handed in. It did not show whether they retained the underlying concepts, could reproduce the quality without AI, detected incorrect suggestions or transferred the reasoning to a new problem later.
That distinction should shape both headlines and classroom decisions.
Critical-thinking training broadened the solution space
The causal-reasoning exercise produced a different pattern. Students who received it explained causes and mechanisms more clearly and generated ideas that were more different from their peers’ ideas.
Those distinctive answers did not earn higher standard rubric scores.
This may be the study’s most important finding. A rubric can be reliable and still reward only a narrow definition of success. If it is calibrated around expected answers, an AI system that produces those answers fluently can perform very well. A student who explores a less conventional option may demonstrate valuable thinking without receiving extra credit for it.
The result does not prove every unusual idea was better. Originality without evidence can be noise. It shows that the evaluation captured conventional quality more readily than difference, which is a problem if independent judgment is one of the intended learning outcomes.
Students did not have to choose between AI and thinking
Students who received both ChatGPT access and causal-reasoning training retained the wider idea variety associated with the lesson while also producing the stronger, more coherent submissions associated with AI access.
OpenAI summarizes the result as complementary effects. That is a reasonable description of this experiment, provided the limits remain visible.
The study does not establish that every AI classroom intervention will complement thinking instruction. It involved one university, one business task, one model and a relatively brief intervention. A mathematics proof, historical interpretation, laboratory report or elementary reading exercise could produce different results.
Still, the experiment suggests a more productive classroom question:
If students can use AI to reach the conventional answer, what must the assignment ask them to do that reveals their own judgment?
Four ways to redesign an assignment
Educators do not need to abandon written work. They can make the reasoning behind it more visible.
1. Ask for assumptions before recommendations
Require students to name the conditions that must be true for their answer to work. In the Bocconi task, a recommendation might assume alumni know where to buy merchandise, care about affiliation or respond to a particular incentive.
An AI can propose assumptions, but the student must decide which are credible and which evidence would test them.
2. Require a rejected alternative
Ask students to present at least two plausible approaches, explain why they rejected one and identify what new evidence could change that choice.
This makes selection part of the assessed work. A polished final answer alone no longer hides the path taken to reach it.
3. Grade failure conditions
Add a rubric item for explaining when the proposal could fail, who might be harmed or excluded and which unintended effect should be monitored.
This directly rewards causal reasoning instead of assuming it will appear inside a general “quality” score.
4. Add a short defense or reflection
Use a two-minute oral explanation, an annotated prompt-and-revision log or a brief reflection identifying what came from the student, what came from AI and what the student changed after verification.
The goal is not surveillance. It is to collect evidence about understanding that the final prose can no longer provide by itself.
For practical boundaries, pair this approach with the parent and teacher guide to ChatGPT for teens, the AI tutors explainer and How to Verify an AI Answer.
What schools should not conclude
Do not use this study to claim that ChatGPT improves critical thinking in every setting. The stronger rubric scores mainly describe the quality of one submitted product.
Do not conclude that critical-thinking instruction is less valuable because it did not raise the conventional grade. The intervention changed idea diversity and causal explanation—outcomes the rubric did not reward.
Do not conclude that access alone is sufficient. Students still need subject knowledge, verification habits and clear expectations about acceptable use.
And do not treat a collaboration announcement as the final word. OpenAI participated in the research and highlighted the results. The randomized design makes the work worth attention, but replication across subjects, ages, institutions and newer models is still necessary.
The takeaway for educators
AI can help a novice produce an answer that resembles expert work. That is useful, but it changes what a finished assignment can prove.
If the goal is professional-quality output, teach students how to use and verify AI. If the goal includes independent judgment, originality and causal understanding, teach those explicitly and make them visible in the assessment.
The strongest classroom design may do both: let students use the tool, then require evidence of the thinking that the tool cannot certify on their behalf.
Frequently asked questions
How many students participated?
Bocconi University reports 1,053 first-year students in economics, finance and management.
Which version of ChatGPT did they use?
Students assigned AI access used ChatGPT Edu with GPT-4o. Results should not be assumed to apply identically to newer or differently configured models.
Did ChatGPT make students more original?
The summaries report that ChatGPT increased the number of ideas but did not drive the wider diversity associated with causal-reasoning training.
Did critical-thinking training improve grades?
Not on the assignment’s standard rubric. It did improve causal explanation and idea distinctiveness, which the rubric did not reward directly.
Should schools allow ChatGPT on every assignment?
No single study can answer that. Permission should depend on the learning goal. If unaided recall or foundational practice is the goal, AI may be inappropriate. If tool-assisted professional work and evaluation are the goal, supervised AI use may be appropriate.
Sources checked
- Bocconi University: Better evaluations with ChatGPT. Critical thinking broadens ideas, published August 27, 2026.
- OpenAI: What students gain from ChatGPT and critical-thinking training, published August 27, 2026.
Sources were checked August 30, 2026. AIViewer did not independently reproduce the experiment or analysis.