What the study found
The study found that generative AI (AI that creates text or other outputs) can be used to score and give feedback on elementary students’ understanding of scientific inquiry with high agreement with human raters. The authors also report that the AI’s feedback was generally appropriate and often supported student agency, although some Korean wording was less clear.
Why the authors say this matters
The authors conclude that GenAI, when properly designed and prompted, can function as a dialogic partner for facilitating students’ epistemic understanding. They also say the study offers new pathways for using GenAI in formative assessment and teacher education, especially for helping students understand the nature of scientific inquiry.
What the researchers tested
The researchers used the Korean version of the Views About Scientific Inquiry for Elementary school students (VASI-E) with 560 responses from 80 fourth-grade students in Korea. They built ChatGPT-4o prompts for scoring and feedback based on established epistemic frameworks, then compared AI scoring with human raters and evaluated the feedback for learner-centered quality.
What worked and what didn't
GenAI scoring showed high agreement with human raters, with overall kappa (a measure of agreement) of 0.825 and item-level values from 0.606 to 0.923. The feedback was rated as generally appropriate, with an average score of 2.75 out of 3, and it was especially strong in promoting student agency through personalized and reflective guidance. Some Korean phrasing created clarity problems.
What to keep in mind
The abstract does not describe longer-term classroom outcomes or whether the system was tested beyond this one group of fourth-grade students in Korea. It also notes language-related challenges in Korean phrasing, which affected clarity.
- GenAI scoring matched human raters closely overall (kappa = 0.825).
- Agreement between AI and humans varied by item, with kappa values from 0.606 to 0.923.
- AI-generated feedback was rated generally appropriate, averaging 2.75 out of 3.
- The feedback was especially strong in supporting student agency.
- Some Korean phrasing in the feedback reduced clarity.