This paper empirically investigates the impact of feedback from Large Language Models (LLMs) on the development of writing proficiency among university-level English learners. Employing a quasi-experimental pre-test/post-test design, the study assigned 52 non-English major freshmen to either an experimental group, which received LLM-generated feedback, or a control group, which received traditional teacher feedback, over a 15-week writing intervention. Writing samples were analyzed quantitatively using the core metrics of Complexity, Accuracy, and Fluency (CAF). The results indicate that the experimental group achieved significantly greater gains in writing accuracy compared to the control group. However, the differences in improvement regarding complexity and fluency between the two groups were not statistically significant. The study concludes that LLM feedback demonstrates a distinct advantage in enhancing the accuracy of linguistic forms, though its effectiveness in promoting linguistic complexity and fluency is less clear within the current model of application. These findings provide concrete empirical support for the use of LLMs in Automated Writing Evaluation (AWE) and offer new perspectives on how L2 writing instruction can effectively harness this technology to realize innovative human-computer collaboration.
Cite this paper
Zhou, Y. (2026). The Impact of LLMs Feedback on L2 Writing Development. Open Access Library Journal, 13, e15998. doi: http://dx.doi.org/10.4236/oalib.1115998.
Zhang, L.J. (2018) Appraising the Role of Written Corrective Feedback in EFL Writing. In: Leung, Y.N., <i>et al</i>., Eds., <i>Reconceptualizing English Language Teaching and Learning in the </i>21<i>st Ce</i><i>ntury</i>, Springer, 161-179.
Bitchener, J. (2012) A Reflection on ‘the Language Learning Potential’ of Written Cf. <i>Journal of Second Language Writing</i>, 21, 348-363. <br>https://doi.org/10.1016/j.jslw.2012.09.006
Kang, E. and Han, Z. (2015) The Efficacy of Written Corrective Feedback in Improving L2 Written Accuracy: A Meta-Analysis. <i>The Modern Language Journal</i>, 99, 1-18. <br>https://doi.org/10.1111/modl.12189
Hyland, K. and Anan, E. (2006) Teachers’ Perceptions of Error: The Effects of First Language and Experience. <i>System</i>, 34, 509-519. <br>https://doi.org/10.1016/j.system.2006.09.001
Grimes, D. and Warschauer, M. (2010) Utility in a Fallible Tool: A Multi-Site Case Study of Automated Writing Evaluation. <i>The Journal of Technology, Learning, and Assessment</i>, 8, No. 6. <br>https://files.eric.ed.gov/fulltext/EJ882522.pdf
Wilson, J. and Shermis, M.D. (2024) Automated Writing Evaluation. In: MacArthur, C.A., Wilson, J., Graham, S. and Allor, J.V., Eds., <i>Handbook of Writing Research </i>(3<i>rd Ed</i><i>ition</i>), The Guilford Press, 257-271.
Gatt, A. and Krahmer, E. (2018) Survey of the State of the Art in Natural Language Generation: Core Tasks, Applications and Evaluation. <i>Journal of Artificial Intelligence Research</i>, 61, 65-170. <br>https://doi.org/10.1613/jair.5477
Biber, D., Nekrasova, T. and Horn, B. (2011) The Effectiveness of Feedback for L1-English and L2-Writing Development: A Meta-Analysis. <i>ETS Research Report Series</i>, 2011, No. 1. <br>https://doi.org/10.1002/j.2333-8504.2011.tb02241.x
Graham, S., Hebert, M. and Harris, K.R. (2015) Formative Assessment and Writing: A Meta-Analysis. <i>The Elementary School Journal</i>, 115, 523-547. <br>https://doi.org/10.1086/681947
Dai, W., Lin, J., Jin, H., Li, T., Tsai, Y., Gašević, D., <i>et al</i>. (2023) Can Large Language Models Provide Feedback to Students? A Case Study on ChatGPT. 2023<i> IEEE International Conference on Advanced Learning Technologies</i> (<i>ICALT</i>), Orem, 10-13 July 2023, 323-325. <br>https://doi.org/10.1109/icalt58122.2023.00100
Mizumoto, A. and Eguchi, M. (2023) Exploring the Potential of Using an AI Language Model for Automated Essay Scoring. <i>Research Methods in Applied Linguistics</i>, 2, Article ID: 100050. <br>https://doi.org/10.1016/j.rmal.2023.100050
Escalante, J., Pack, A. and Barrett, A. (2023) AI-Generated Feedback on Writing: Insights into Efficacy and ENL Student Preference. <i>International Journal of Educational Technology in Higher Education</i>, 20, Article No. 57. <br>https://doi.org/10.1186/s41239-023-00425-2
Hunt, K.W. (1970) Syntactic Maturity in Schoolchildren and Adults. <i>Monographs of the Society for Research in Child Development</i>, 35, iii-iv, 1-67. <br>https://doi.org/10.2307/1165818
Foster, P. and Wigglesworth, G. (2016) Capturing Accuracy in Second Language Performance: The Case for a Weighted Clause Ratio. <i>Annual Review of Applied Linguistics</i>, 36, 98-116. <br>https://doi.org/10.1017/s0267190515000082