In a randomized trial with more than 1,000 Bocconi University first-years, ChatGPT access was worth nearly a full point on a five-point rubric, while the originality gains from causal-reasoning training showed up only in automated text analysis.
Students who could use ChatGPT handed in work that was better in quality and hung together more coherently, according to OpenAI's announcement; a causal-reasoning task — one type of critical thinking — instead led students to produce more unique ideas. The students who got both registered both effects.
Researchers at Bocconi University ran the randomized trial with OpenAI Economic Research, and OpenAI published the findings on its own site — research favourable to its own product, co-authored and released by the company that sells it.
More than 1,000 first-year undergraduates took on a real business problem: marketing proposals for the university's merchandise shop. Class periods were randomly assigned to one of four arms — use of ChatGPT, instruction in causal reasoning, the two combined, or neither.
Causal reasoning here means tying causes to effects and setting out the reasons a proposed solution might work or might not. The instruction had nothing to do with AI.
Trained human graders scored the submissions against a five-point rubric. Separately, automated text analysis gauged how many ideas each piece held and how varied they were, whether it carried traces of causal reasoning, and how closely it resembled work produced by three experts.
On that five-point scale, students with ChatGPT available to them came in nearly a whole point ahead. What they turned in carried more ideas, ran on tighter logic and looked closer to the experts' recommendations.
Students were not simply handing the assignment to ChatGPT, the researchers say: working out what to ask it, judging what came back and picking which parts made it into the final version stayed with them.
The instruction arm looked different. Those students set out more clearly why their ideas could succeed and the circumstances in which they could fall short, yet their rubric marks were no higher, because the rubric gauged nothing beyond how well a proposal served two conventional marketing aims: lifting awareness of the store and getting more people to use it.
The text analysis caught what the marks did not — taken as a whole, the arm produced a broader spread of ideas that stood further apart from what other students came up with. A conventional rubric can credit a lucid, well-organized answer and still miss the fact that a student arrived at an idea nobody else had.
Students handed both showed the variety of the activity-only arm alongside the rubric marks and idea counts of the ChatGPT-only arm. Their work carried more signs of searching for explanations and of challenging assumptions, and improved on a broader set of measures than any other group's.
Where AI helps students turn out finished, expert-seeming work, judging by the handed-in answer alone reveals less about what a student genuinely grasps. Coursework may have to change so that it can gauge other attributes, originality among them.