A randomized experiment with over 1,000 Bocconi University students found that access to ChatGPT and training in causal reasoning improve student work in distinct, complementary ways. Students using ChatGPT produced higher-quality, more coherent answers similar to expert recommendations, while those receiving causal reasoning training generated a wider variety of unique ideas.
- Access to GPT-4o improved five-point rubric scores by almost a full point and increased logical coherence.
- Causal reasoning exercises led students to explain why ideas might work or fail and produced more distinct ideas not captured by standard grading.
- Students receiving both interventions showed gains across the widest range of measures, combining polished output with originality.
The results suggest that as AI makes it easier to produce polished answers, educational assessments must adapt to measure other qualities like originality and reasoning.