“Jeder hat seinen eigenen Geschmack”: Comparative Analysis of AI and Human Raters in German Proverb Oral Performance Assessment
DOI:
https://doi.org/10.63011/ip.v3i3.88Keywords:
Artificial Intelligence, language assessment, human rater, oral performance, German proverbsAbstract
This study aims to analyze scoring variations and test the statistical significance of differences between AI-based evaluation (Large Language Model) and human raters in assessing the oral performance of German proverbs. Using a quantitative, within-subjects design, the study involved 30 students who were evaluated on four parameters: grammar, pronunciation, fluency, and the use of proverbs. The results of a paired-samples t-test reveal highly significant differences (p < 0.001) across all assessment aspects. The AI consistently demonstrates a leniency bias, assigning higher absolute scores than human raters. The largest discrepancy is in the use of proverbs, with a mean difference of 3.73 points. These findings indicate that AI tends to operate at the level of structured linguistic surface features, yet remains limited in capturing cultural nuances, implicit meanings, and pragmatic contextual appropriateness. These are areas that constitute the core strength of human cognitive sensitivity. This study recommends implementing a hybrid app using language to evaluate, in which AI assesses structural-mechanistic aspects, while interpretive sociolinguistic dimensions remain under human raters' control.
References
Aliyeva, E. (2025). The Role of Teaching Proverbs and Sayings in Enhancing Students' Speaking Skills. Acta Globalis Humanitatis et Linguarum, 2(1), 54-61. https://doi.org/10.69760/aghel.02500107
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021, March). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency (pp. 610-623). https://doi.org/10.1145/3442188.3445922.
Charteris-Black, J. (1995). Proverbs in communication. Journal of Multilingual & Multicultural Development, 16(4), 259-268.
https://www.google.com/search?q=https://doi.org/10.1080/01434632.1995.9994609 Cram, D. (2015). The linguistic status of the proverb. In Wise Words (RLE Folklore) (pp. 73-97). Routledge. https://www.google.com/search?q=https://doi.org/10.4324/9781315662138-11
Ercikan, K., & McCaffrey, D. F. (2022). Optimizing implementation of artificial-intelligence-based automated scoring: An evidence centered design approach for designing assessments for AI-based scoring. Journal of Educational Measurement, 59(3), 272-287. https://doi.org/10.1111/jedm.12332
Hartwell, K., & Aull, L. (2023). Editorial Introduction–AI, corpora, and future directions for writing assessment. Assessing Writing, 57, 100769. https://doi.org/10.1016/j.asw.2023.100769
Li, B., Qunhan, X., & Mao, C. (2026). Differences between human and AI scoring: A meta-analysis of English language assessments. Scientific Reports. https://doi.org/10.1038/s41598-026-48053-w
Lundgren, M. (2024). Large language models in student assessment: Comparing ChatGPT and human graders. arXiv preprint arXiv:2406.16510. https://doi.org/10.48550/arXiv.2406.16510
Mahowald, K., Ivanova, A. A., Blank, I. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko,
E. (2024). Dissociating language and thought in large language models. Trends in cognitive sciences, 28(6), 517-540.
https://www.sciencedirect.com/science/article/pii/S1364661324000275
Margetson, K., McLeod, S., Verdon, S., & Tran, V. H. (2023). Transcribing multilingual children’s and adults’ speech. Clinical Linguistics & Phonetics, 37(4-6), 415-435. https://doi.org/10.1080/02699206.2022.2051073
Mieder, W. (2004). Proverbs: A Handbook. Greenwood Press.
Ranalli, J., Link, S., & Chukharev-Hudilainen, E. (2017). Automated writing evaluation for formative assessment of second language writing: Investigating the accuracy and usefulness of Grammarly. CALICO Journal, 34(2), 149-181.
https://doi.org/10.1080/01443410.2015.1136407
Rietveld, T., & van Hout, R. (2017). The paired t test and beyond: Recommendations for testing the central tendencies of two paired samples in research on speech, language and hearing pathology. Journal of communication disorders, 69, 44-57.
Yanga, T. (2017). Inside the proverbs: A sociolinguistic approach. In African Languages/Langues Africaines (pp. 130-157). Routledge.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Sudarmaji, Iman Santoso, Aditya Rikfanto, Retna Endah Sri Mulyati

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.




