The UK government is considering a cost-cutting measure that would fundamentally alter how writing assessments are moderated in schools. Rather than use authentic samples from real children, moderators evaluating key stage 2 literacy standards would judge fairness against artificial writing generated by ChatGPT.
The proposal, reported by New Scientist, reflects budget pressures on England's education system. Key stage 2 assessments test children aged 10 and 11 across English and mathematics. Moderators currently receive real examples of student work to calibrate their grading standards and ensure consistency across schools. Replacing these authentic samples with AI-generated text would reduce administrative costs significantly.
The scheme raises immediate pedagogical and validity concerns. AI language models like ChatGPT produce text that differs structurally from how children actually write. Children make characteristic errors, show developmental patterns in syntax and spelling, and demonstrate emerging understanding of conventions. Their authentic mistakes reveal thinking processes. ChatGPT generates polished, statistically average text without the idiosyncratic features that mark genuine developmental writing.
Moderators trained on AI-generated examples might develop skewed expectations about what children at each attainment level should produce. They may undervalue creative variation or penalize authentic developmental errors that appear in real student work but not in sanitised AI output. This misalignment between training examples and actual student writing could distort grade boundaries upward, inflating achievement metrics while leaving genuine differences in student capability undetected.
The plan also raises questions about what ChatGPT actually represents pedagogically. Language models trained on internet text capture statistical patterns rather than developmental progression. A child learning to write progresses through predictable stages: simple sentences, then compound sentences, then embedded clauses. They acquire spelling conventions gradually. These stages reflect cognitive development. ChatGPT has no developmental trajectory. It mirrors patterns from its training data, which skews toward published, edited writing rather than authentic student expression.
There are also practical risks. ChatGPT can produce plausible-sounding but technically incorrect grammar or factual claims. If moderators use these examples to establish standards, they might inadvertently normalise errors into the grading framework. This contamination of standards spreads downstream, affecting how teachers understand expectations and how students learn.
The government department responsible for education assessment has not formally confirmed the proposal as policy. Officials may be exploring it as one option among several cost-reduction strategies. However, the plan reveals the pressure exerted on public assessment systems by budget constraints. Quality moderation requires human expertise and authentic materials, both expensive. Automation and AI substitution offer apparent shortcuts.
Education bodies internationally face similar pressures. Finland and Singapore have scaled assessments using digital platforms, but they retained authentic student work in moderation processes. None has yet replaced human exemplars with AI-generated text as the primary standard-setting tool.
The proposal warrants scrutiny before implementation. Cost savings in assessment moderation risk downstream costs in educational validity and student outcomes.
