Do student ratings match teacher ratings? Comparing rater severity and reliability in L2 writing assessment.
- Publication History
- Published online: October 31, 2026
- DOI
- https://doi.org/10.35877/454RI.daengku4981
- Copyright
- Copyright (c) 2026 Nurul Fariena Asli, Mohd Effendi Ewan Mohd Matore; Nurul Iman Ahmad Bukhari
- User License
- https://creativecommons.org/licenses/by-nc-sa/4.0
Abstract
This study examines rater severity and rater reliability in a writing assessment across teacher and student raters within the Malaysian secondary schools L2 context. Four Primary Trait Writing (PTW) rubrics were developed and validated to be utilized in this self-assessment activity. A quantitative research approach was employed, involving 149 secondary school students and three English language teachers from six public schools in Malaysia. Teacher assessment and student self-assessment were used to evaluate students’ essays, and the Many Facet Rasch Model (MFRM) model was applied using FACETS to calibrate and analyse the scores. The analysis focused on identifying differences in rater severity and examining the consistency of rubric application between teacher and student raters. Results showed that teacher and student raters exhibited similar severity levels, with no significant differences between the two groups (?² = 0.0, df = 1, p = .90). Furthermore, acceptable fit statistics (MnSq = 0.5-1.5; Zstd within ±2) indicated consistent application of the PTW rubrics across both rater groups. These findings suggest that both teacher and student raters shared a common understanding of the assessment criteria and functioned as a homogeneous rater group, demonstrating similar levels of severity and reliability in the assessment of L2 writing performance.
Keywords
Citation
Statements

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.