“Jeder hat seinen eigenen Geschmack”: Comparative Analysis of AI and Human Raters in German Proverb Oral Performance Assessment

Authors

  • Sudarmaji Universitas Negeri Yogyakarta
  • Iman Santoso Universitas Negeri Yogyakarta
  • Aditya Rikfanto Universitas Negeri Yogyakarta
  • Retna Endah Sri Mulyati

DOI:

https://doi.org/10.63011/ip.v3i3.88

Keywords:

Artificial Intelligence, language assessment, human rater, oral performance, German proverbs

Abstract

This study aims to analyze scoring variations and test the statistical significance of differences between AI-based evaluation (Large Language Model) and human raters in assessing the oral performance of German proverbs. Using a quantitative, within-subjects design, the study involved 30 students who were evaluated on four parameters: grammar, pronunciation, fluency, and the use of proverbs. The results of a paired-samples t-test reveal highly significant differences (p < 0.001) across all assessment aspects. The AI consistently demonstrates a leniency bias, assigning higher absolute scores than human raters. The largest discrepancy is in the use of proverbs, with a mean difference of 3.73 points. These findings indicate that AI tends to operate at the level of structured linguistic surface features, yet remains limited in capturing cultural nuances, implicit meanings, and pragmatic contextual appropriateness. These are areas that constitute the core strength of human cognitive sensitivity. This study recommends implementing a hybrid app using language to evaluate, in which AI assesses structural-mechanistic aspects, while interpretive sociolinguistic dimensions remain under human raters' control.

References

Aliyeva, E. (2025). The Role of Teaching Proverbs and Sayings in Enhancing Students' Speaking Skills. Acta Globalis Humanitatis et Linguarum, 2(1), 54-61. https://doi.org/10.69760/aghel.02500107

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021, March). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency (pp. 610-623). https://doi.org/10.1145/3442188.3445922.

Charteris-Black, J. (1995). Proverbs in communication. Journal of Multilingual & Multicultural Development, 16(4), 259-268.

https://www.google.com/search?q=https://doi.org/10.1080/01434632.1995.9994609 Cram, D. (2015). The linguistic status of the proverb. In Wise Words (RLE Folklore) (pp. 73-97). Routledge. https://www.google.com/search?q=https://doi.org/10.4324/9781315662138-11

Ercikan, K., & McCaffrey, D. F. (2022). Optimizing implementation of artificial-intelligence-based automated scoring: An evidence centered design approach for designing assessments for AI-based scoring. Journal of Educational Measurement, 59(3), 272-287. https://doi.org/10.1111/jedm.12332

Hartwell, K., & Aull, L. (2023). Editorial Introduction–AI, corpora, and future directions for writing assessment. Assessing Writing, 57, 100769. https://doi.org/10.1016/j.asw.2023.100769

Li, B., Qunhan, X., & Mao, C. (2026). Differences between human and AI scoring: A meta-analysis of English language assessments. Scientific Reports. https://doi.org/10.1038/s41598-026-48053-w

Lundgren, M. (2024). Large language models in student assessment: Comparing ChatGPT and human graders. arXiv preprint arXiv:2406.16510. https://doi.org/10.48550/arXiv.2406.16510

Mahowald, K., Ivanova, A. A., Blank, I. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko,

E. (2024). Dissociating language and thought in large language models. Trends in cognitive sciences, 28(6), 517-540.

https://www.sciencedirect.com/science/article/pii/S1364661324000275

Margetson, K., McLeod, S., Verdon, S., & Tran, V. H. (2023). Transcribing multilingual children’s and adults’ speech. Clinical Linguistics & Phonetics, 37(4-6), 415-435. https://doi.org/10.1080/02699206.2022.2051073

Mieder, W. (2004). Proverbs: A Handbook. Greenwood Press.

Ranalli, J., Link, S., & Chukharev-Hudilainen, E. (2017). Automated writing evaluation for formative assessment of second language writing: Investigating the accuracy and usefulness of Grammarly. CALICO Journal, 34(2), 149-181.

https://doi.org/10.1080/01443410.2015.1136407

Rietveld, T., & van Hout, R. (2017). The paired t test and beyond: Recommendations for testing the central tendencies of two paired samples in research on speech, language and hearing pathology. Journal of communication disorders, 69, 44-57.

Yanga, T. (2017). Inside the proverbs: A sociolinguistic approach. In African Languages/Langues Africaines (pp. 130-157). Routledge.

Downloads

Published

10.08.2026