|
Vol.15, No.3, August 2026. ISSN: 2217-8309 eISSN: 2217-8333
TEM Journal
TECHNOLOGY, EDUCATION, MANAGEMENT, INFORMATICS Association for Information Communication Technology Education and Science |
AFAR: Automated Feedback for Arabic Responses of Short-Answer Questions
Sara Alqaidi, Omaima Almatrafi, Arwa Wali
© 2026 Sara Alqaidi, published by UIKTEN. This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License. (CC BY-NC-ND 4.0)
Citation Information: TEM Journal. Volume 15, Issue 3, Pages 2772-2784, ISSN 2217-8309, DOI: 10.18421/TEM153-63, August 2026.
Received: 26 August 2025.
Abstract:
Providing feedback is a vital component of the effective learning process. It guides students in recognizing their weaknesses and addressing areas for improvement, rather than merely assigning final scores. However, providing manual feedback is often time-consuming and labour-intensive for instructors, especially when managing a large number of students. Previous studies have explored various approaches for automatic feedback generation, including rule-based systems and machine learning techniques. More recently, research has investigated the use of large language models (LLMs), such as GPT, to automate feedback generation. However, most of these works have focused on English-language responses and technical subjects. There is still limited research applying LLMs to provide feedback for short-answer tasks in low-resource languages, such as Arabic. To bridge this gap, this study introduces AFAR (Automated Feedback for Arabic Responses), a novel framework designed to generate personalized feedback on students’ short-answer responses in Arabic using GPT-4. AFAR comprises prompt engineering and few-shot learning techniques to guide GPT-4 in producing meaningful feedback. To ensure quality, the feedback was assessed by two human evaluators in terms of coherence, relevance, clarity, and personalization. The results showed that the feedback generated by the GPT-4 was generally high in quality, with 90% rated as coherent, 82% as relevant, 92% as clear, and 85% as personalized.
Keywords – Prompt-based learning, Generative artificial intelligence, feedback quality, evaluation. |
|
----------------------------------------------------------------------------------------------------------- ----------------------------------------------------------------------------------------------------------- |