?
Correcting or Rewriting? An Expert Evaluation of LLM-Based GEC on Academic Learner Data
This paper investigates how large language models correct complex grammatical errors in Russian academic
learner writing. Unlike traditional minimal-edit GEC systems, LLMs often apply generative rewriting strategies that
may improve fluency, but risk structural overcorrection and semantic drift. We introduce a new expert benchmark
derived from an authentic 3,1M-word learner corpus and construct an evaluation set annotated for error type and
complexity.
We propose an expert-driven evaluation framework combining quantitative scoring, structural-change analysis,
and blind pairwise comparison. Results reveal a consistent minimal-edit vs. generative trade-off across LLMs. This
trade-off has direct implications for evaluation, as purely reference-based metrics may underrepresent structural
overcorrection and fail to capture differences in correction strategies.