AI Grades Essays Higher Than Human Teachers Do — and That Should Worry Anyone Hoping It Would Fix Grade Inflation
CARDIFF/CAMBRIDGE — Two separate research teams published findings this year that arrive at the same uncomfortable conclusion from different angles: when AI grades student essays, it doesn't just disagree with human teachers — it disagrees in a specific, predictable direction. It grades higher. Consistently, and in some cases dramatically.
Researchers at Cardiff University and the University of Melbourne uploaded 50 undergraduate bioscience essays to two versions of ChatGPT, scoring each one against seven assessment criteria under four different prompting setups, then checked every AI-generated score against the mark a human academic had already assigned the same essay. In nearly every condition tested, the AI models returned higher averages than their human counterparts — and in one case, the gap between an AI grade and a human grade hit 40 points on a 100-point scale. The pattern wasn't random noise, either. Lower-scoring essays saw the biggest jump in AI-assigned marks, while some of the strongest essays actually got marked down relative to what a human examiner had given them. The study, published in Assessment & Evaluation in Higher Education, concluded the large language models tested "varied considerably" in the marks they awarded and were, in the researchers' words, "inadequate predictors of the human mark awarded to essays."
A separate, larger Cambridge-led study — testing three frontier AI systems, including the latest 2026 versions of Claude and ChatGPT, on more than 750 real student essays from three UK universities — found something that should worry anyone hoping AI grading might at least be consistently generous, even if inflated. The AI systems only matched the broad degree classification a human examiner had given — a first, a 2:1, a 2:2, and so on — somewhere between 35% and 65% of the time, depending on the system and essay type. And the errors weren't evenly distributed: AI models routinely undervalued work that human examiners had ranked among the very best, while overvaluing essays that humans had ranked among the weakest — the opposite of what you'd want from a grading tool, which is to reliably distinguish excellent work from mediocre work at exactly the extremes where that distinction matters most.
The researchers identified what's actually driving the inflation, and it's a genuinely revealing finding: AI systems were, in the study's own phrasing, "oversensitive to linguistic features" — handing out higher marks based on essay length, vocabulary range, and sentence complexity, regardless of whether the underlying academic argument was actually any good. In plain terms, the AI wasn't grading whether a student understood the material. It was grading whether the essay sounded like it was written by someone who understood the material — rewarding style over substance, exactly the failure mode Cambridge's own summary of the research warned about.
This lands at a genuinely awkward moment for anyone hoping AI might be part of the solution to grade inflation rather than an accelerant of it. As covered previously on this site, Harvard just voted to cap the share of A's any single course can award, after finding solid A's had climbed to 60% of every undergraduate grade — up from roughly a quarter two decades earlier. The uncomfortable question these new studies raise: if universities lean on AI to help manage grading workloads as writing assignments scale up, and AI systems demonstrably grade more generously than human examiners while rewarding surface-level linguistic polish over genuine understanding, AI-assisted grading doesn't fix grade inflation. It risks quietly making it worse, at exactly the moment institutions are trying to reverse it.
There's a human dimension to this too, and it's not a minor footnote. Student focus groups run as part of the Cambridge research found some participants said they would feel "cheated" if their coursework were assessed primarily by AI rather than a human academic — a consent question UK universities considering AI marking tools will now have to navigate alongside the accuracy question, particularly in fee-paying programs where students are explicitly paying for expert human engagement with their work. Separately, data cited in reporting on the UK's own House of Lords briefing on AI found two-thirds of secondary school teachers already believe their pupils' critical thinking has measurably declined under AI usage — a finding that echoes, almost exactly, the "cognitive offloading" concerns raised in Brookings' recent research on AI in classrooms.
Not every study paints quite as bleak a picture. Research from UC Irvine comparing ChatGPT against expert human evaluators on 1,800 middle and high school essays found "fair" to "moderate" agreement overall, with ChatGPT landing within one point of a trained human grader roughly 76% to 89% of the time depending on subject and essay batch — meaningfully closer than the Cardiff and Cambridge findings suggest, though still with real, direct-to-teacher-workflow risk built in for the essays where it disagreed. The gap between these findings likely comes down to what's being measured and how: shorter, more structured assignments seem to narrow the disagreement; longer, more open-ended university-level essays widen it.
The Cambridge team's own conclusion is likely to become the default policy position across UK universities through the rest of 2026: a human should always determine the final mark, with AI restricted to first-pass screening or feedback drafting rather than autonomous grading. For education systems elsewhere — Azerbaijan included, as it continues weighing exactly how deeply to integrate AI into its own classrooms rather than making it a standalone subject — the lesson from this research runs parallel to nearly every other AI-in-education finding covered on this site this year: the technology isn't simply good or bad at the task. It's a tool whose failure modes are specific, measurable, and dangerous precisely when nobody's watching closely enough to catch them — which, for a system already struggling with grade inflation, is exactly the wrong moment to hand over more of the grading pen.