My Research in the News: Why Does AI Struggle with Korean CSAT Earth Science?

The joint research recently conducted by Seoul National University and National Taiwan Normal University (NTNU) was featured in an article by ‘AI Matters’. Having traveled back and forth between Taiwan and “Korea” pondering over science education research, it feels deeply refreshing to see it introduced in such a public-facing article. I’d like to briefly share with my blog neighbors what this research is about and what educational significance it holds.

What did we research?

This study analyzes how the latest Large Language Models (LLMs), such as GPT-4o, Gemini 2.5 Flash, and Gemini 2.5 Pro, solve the 2025 CSAT (“Korea”’s College Scholastic Ability Test) ‘Earth Science I’ questions. Beyond simply measuring “what score the AI received,” we focused on the fundamental question: “How does AI fail differently from humans, and why?” We gave the AI the entire test paper as is, provided it question by question, and even strictly separated the text from the diagrams. Under optimal conditions, the Gemini 2.5 Pro model scored 68, approaching the performance of top-tier students. However, what our research team focused on was the highly unique and repetitive patterns that emerged when the AI failed.

How does AI fail?

The analysis revealed that AI exhibits three fatal error patterns completely different from how humans make mistakes.

  • Perception-Cognition Gap: A phenomenon where AI perceives the visual information itself but fails to read the scientific rules contained within it. For instance, while it could see a radial graph showing the wind direction of a typhoon, it could not connect the lines to their scientific meanings of ‘clockwise’ and ‘counterclockwise’.
  • Calculation-Conceptualization Discrepancy: Cases where the mathematical calculation is performed perfectly, but the physical and scientific concepts the numbers represent are not grasped. In other words, procedural calculation ability and conceptual understanding are completely disjointed.
  • Process Hallucination: A phenomenon where, in a situation requiring step-by-step reasoning based on complex data, the AI simply skips the process and fits superficially related background knowledge to jump to a conclusion.

What is the meaning of this research?

The true value of this research does not lie in merely proving that “AI has not yet surpassed humans.” The key point is that it provides a clear clue as to how our educational assessment methods must change in the approaching AI era. This is exactly why the design of ‘AI-resistant assessment’, which I also presented at the “Korean” Association for Science Education (KASE) winter conference last February, is so crucial. Assessments asking for simple knowledge or procedural calculations that AI can easily solve will no longer hold discriminatory power. Instead, we need assessments that ask students to interpret hidden rules in unstructured diagrams or to apply the conceptual meaning of calculated results to real-world situations. This will become the direction for effective educational assessment that measures students’ true understanding while being difficult for AI to easily mimic.

Moving forward, I will continue to conduct meaningful research that incorporates the perspective of learning science at the intersection of technology and education.

You can check the detailed article content and the original report published on arXiv via the links below. Since it was uploaded to arXiv as a preprint, the writing is not yet perfectly polished. I need to submit it to a formal academic journal soon, but given the rapid pace at which AI changes, it is not easy to find the time to do so.

👉 View the original article: ChatGPT, which passed the medical exam, has a mental breakdown in front of CSAT Earth Science… – AI Matters

👉 Paper (arXiv): “ChatGPT and Gemini participated in the “Korean” College Scholastic Ability Test – Earth Science I”

About: Seok-Hyun Ga

I am currently a Senior Researcher at the Center for Educational Research at Seoul National University, having previously served as a Visiting Assistant Professor at the Graduate Institute of Science Education and Science Education Center at National Taiwan Normal University. I hold a Ph.D. and an M.S. in Science Education, as well as a B.S. in Earth Science Education and Physics Education from Seoul National University. Additionally, I earned a B.S. in Computer Science from Korea National Open University.


One thought on “My Research in the News: Why Does AI Struggle with Korean CSAT Earth Science?”

Leave a Reply

Your email address will not be published. Required fields are marked *