When Less is More: AI Feedback for Language Learners 

For this post, let me jump straight in by asking you your opinion. Imagine you’re learning English, and you ask ChatGPT to correct an English sentence you have written. You tell it, “Correct my English: I enjoy to watch movies”.  
 
You try the same prompt on two different chats, and get two different responses. Which response, (a) or (b) below, do you think would be more useful to you as someone learning English? 
 
(a) 

 
(b) 

Which one do you think would be more useful for an English learner? (a) or (b)?  

In recent years while teaching English in South Korea, I’ve watched a lot of learners discover AI writing tools. A student writes a diary entry or an essay, pastes it into ChatGPT, asks it to “correct my English”, then gets back a polished, fluent, native-sounding version, just like (a) above. On the surface, this looks wonderful. However, as an English teacher, I have learnt through trial and error that the most effective feedback is more like the simple response in (b).  
 
The reason (a) isn’t as useful is that the LLM has over-corrected the English. Rather than gently fixing the actual mistake, LLMs tend to rewrite the whole sentence into something more idiomatic and “native-like”. In doing so, they quietly erase the learner’s own voice, bury the original error under a pile of stylistic changes and often introduce vocabulary and structures well beyond the learner’s current level. The student ends up with a lovely paragraph they didn’t write and can’t fully understand, and crucially, they never actually see what they got wrong. 

This matters because of a well-established idea in language learning called the Noticing Hypothesis (Schmidt, 2001): learners tend to acquire a form when they consciously notice it. A teacher’s small, surgical correction in a piece of writing helps them notice. A wholesale AI rewrite hides the very thing they needed to see. So a colleague and I recently wrote a paper asking “Can we get a large language model (“LLM” or simply “model”) like GPT-4o to hold back – to make minimal, targeted edits rather than overwhelming a learner with a full rewrite? And what’s the best way to do this?” 

So what did we find? 

We compared two broad strategies for controlling the LLM’s behaviour, using a set of 1,000 short English essays written by Korean people. 

The first strategy was prompting – simply telling the model how to behave, using carefully written instructions like “only fix clear errors, don’t swap words for synonyms, preserve the learner’s meaning” and so on.  
 
The second was fine-tuning – actually retraining the model on examples of the minimal-edit style we wanted. 

The result was clear, and an eye-opener for anyone who spends hours perfecting their prompts: fine-tuning won, decisively. No matter how cleverly we worded our instructions, the off-the-shelf LLM kept drifting back to its old habit of over-correcting. Fine-tuning, by contrast, “baked in” the restraint. The behaviour became stable and consistent instead of something we had to keep wrestling the model into on every single run. 

The second finding was the one I find most encouraging for teachers. We expected that fine-tuning would require a huge dataset of many thousands of examples to display consistent, useful behaviour. It didn’t. A relatively small, high-quality set of around 500 examples delivered most of the benefit. Bigger datasets nudged precision and fluency up a little further, but the big leap came early. In other words, quality and targeting mattered far more than sheer quantity. 

Why this matters for teachers 

I think the real significance of the paper isn’t really about grammar correction at all. It’s about who gets to decide how an AI tool behaves. 

Since ChatGPT took the world by storm in late 2022, the assumption has been that shaping how an AI behaves is a job for huge technology companies with vast datasets and deep pockets. Our findings quietly challenge that. If a few hundred carefully corrected examples can retrain a model to give exactly the kind of feedback you want, then the person best placed to build the tool is the teacher who actually knows the learners. A teacher, or a small department, can take a general-purpose LLM and bend it towards their own students, their own level, their own priorities. Fine-tuning, in other words, hands the controls back to the people doing the teaching. I find that a genuinely encouraging thought. 

From one paper to a whole corpus: my PhD research 

This paper was one of the seeds of my PhD research. The fine-tuning in the study drew on a database of essays written by Korean learners of English. It was made up of written data only and grammar-focused. It showed me two things: that targeted data can reliably steer an AI’s behaviour, and that generic tools systematically miss the specific, patterned errors that Korean learners of English make. Anyone who has taught Korean learners will recognise the struggle with the English articles “a”, “an” and “the” – simply because Korean has no direct equivalent. A generic model treats these as random mistakes, but they aren’t random at all. They’re systematic, and they’re a fingerprint of the learner’s first language. 

My PhD sets out to build the resources that will let us capture patterns like these properly and then use them to help people learn more effectively: a multimodal (writing and speaking), longitudinal (samples spaced over time) corpus of Korean-English interlanguage (a learner’s developing, rule-governed grammar system), followed by an AI teaching system built on this corpus. Where the paper used written essays, the PhD adds spoken data too, collected in waves over time so we can watch learners’ English develop. Data collection is already well underway – on top of around 50,000 short English essays collected since 2017, I’ve recently interviewed dozens of participants and have been building a transcription pipeline to turn all that speech into a properly annotated, searchable corpus. 

There’s an irony here worth dwelling on. As large language models increasingly flood the written world with machine-generated text (contributing to the “Dead Internet Theory”), authentic spoken language – real people talking, hesitating, code-switching and making genuine errors – is becoming more precious, not less. A carefully collected spoken learner corpus is exactly the kind of trustworthy, human-grounded resource that this AI-saturated moment calls for. 

Where CASS comes in 

All of this sits within the tradition CASS has helped build. The centre’s work includes foundational corpora like the British National Corpus 2014 and the Trinity Lancaster Corpus. My own corpus (working title: Korean Multimodal English Corpus – K-MEC) will be a small, specialised cousin of these, deliberately built in line with the Lancaster standard so that it can be used by corpus linguistics researchers. 

CASS is also where the linguistic and the computational meet in exactly the way my research needs. My supervisor, CASS Co-Director Vaclav Brezina, has written about how corpus linguistics can respond to the rise of LLMs and the “black box” problem – the very issue that motivates my attempt to build transparent, explainable, L1-sensitive AI feedback rather than another opaque tool. And practical instruments like #LancsBox X are what make a corpus like mine genuinely usable, both for me and for other researchers. 

More broadly, CASS’s whole mission of bringing the corpus approach to real human and social questions, from healthcare communication to language learning, is a reminder that the point of all this data is people. My paper is fundamentally an argument that AI in education should serve learners on their own terms: preserving their voice, meeting them at their level and helping them notice and grow. That’s a very corpus-linguistic, and a very CASS, way of thinking about technology. 

Looking ahead 

The paper answered a narrow question: how do you get an AI to make minimal, learner-friendly grammar corrections? The answer was to “fine-tune it on a small amount of good, targeted data.” My PhD now asks a bigger question: what if that targeted data were a wide-ranging, multimodal record of how Korean learners develop both their spoken and written English over time? Through this we can have not just another learner corpus, but a foundation for better understanding a whole learner population – and a blueprint for using AI in education responsibly. 

If you would like to read the paper, it is open access and can be found here (for the full text, click the “download” button on the page):  
Sumner, J. P., & Dillon, T. (2026). Evaluating Prompting and Fine-Tuning Strategies for GPT-Based Grammatical Error Correction. Language Education & Assessment, 9, 103750.  
https://doi.org/10.29140/lea.2026.103750 
 
References: 
Berghel, H. (2026). Generative AI is Breathing New Life Into the Dead Internet Theory. Computer, 59(1), 132–139.  
https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=11320999 
Schmidt, R. (2001). Attention. In P. Robinson (Ed.), Cognition and second language instruction (pp. 3–32). Cambridge University Press.  
https://nflrc.hawaii.edu/PDFs/SCHMIDT%20attention.pdf 

Thanks for reading. If you would like to discuss my research, please email me at j.p.sumner@lancaster.ac.uk