Further thoughts on AI in assessments
July 26, 2026•1,010 words
Discussing the recently expressed concerns about students using LLMs to 'cheat' with academic colleagues, I am surprised how many think that we should return to closed exams.
Setting aside the inherently discriminatory nature of that form of assessment, this is deeply ironic. Since confidence is a necessary condition of doing well in exams, that mode of assessment rewards the confident-plausible-but-superficial-and-wrong answer (the kind of thing an LLM might produce) over the unconfident-cautious-nuanced-subtle-but-built-on-deep-understanding answer (the kind of thing only a human can produce).
However, the irony masks a much bigger problem with the call for a return to closed exams. The premise of this call is:
We cannot tell the difference between essays produced by students we have taught and those produced by LLMs.
Thus we have to actually watch the students writing to know the essay is human-produced.
But that premise is exactly what the AI companies want everyone to believe. They are aiming to produce a product which will replace the labour currently done by graduates. The promise of that is what 'justifies' their ridiculous stock market valuations and all the rhetoric about the 'fourth industrial revolution'.
Universities current produce the labour force for graduate jobs. If we go into a panic about not being able to tell the difference between what a graduate produces and what an LLM produces, then we are loudly signalling to the businesses that currently employ graduates that LLMs could do the job just as well. And thereby they reinforce the inevitability narrative, maintain the AI bubble, and undermine a core social function of higher education (and thus the basis of their funding).
This is a general problem with moral panics: they are exploited to legitimise social changes which increase the power of those who already have it at the expense of those who don't.
What instead?
I have argued elsewhere that I do not think there is a problem with 'AI cheating' and there are things universities, and primarily their leaders, could do which would reduce the incidence of students using LLMs when we don't want them to. We shouldn't be trying to find 'AI proof' assessments but instead addressing the reasons why students - who by and large do not want to cheat - are using LLMs.
However, that does not address the premise which is driving this ill-advised call for a return to exams:
We cannot tell the difference between essays produced by students we have taught and those produced by LLMs.
Of course, the panic is driven in part by reading essays which feel like they are produced by, or with a lot of assistance from, LLMs. That feeling comes from years of experience reading student work, which is stylistically very different from the fluent, grammatical and well-structured LLM output. But you cannot mark an essay down for being 'well-written'.
Instead, we should reflect on why student writing is like it is and ask whether that reflects something of value the student is doing which the LLM is not.
Student writing
A genuine student essay is written as part of the process of gaining understanding. Every student continues their learning and deepens their understanding of key concepts through the process of expressing ideas, structuring arguments, deciding what is and is not relevant etc. A student essay is an edited stream of consciousness as the mind works through the question. It thus retains much of the messiness of the steam of consciousness.
Academic research - including PhDs - involves a further stage of rewriting that into a different genre which is no longer a stream of consciousness, but instead a closer approximation to the Platonic form of the answer to the research question. LLMs are trained to produce output which imitates this second genre. Even when in 'Reasoning Mode' there is a clear sense in which the output is an idealisation of real thought processes.
This is why LLM output is so different from student work and is also so depressing to read: as educators what we want to see in an essay are the actual thought processes, we want to see the struggle of thinking something through. But why?
The University 'Value Proposition'
The distinction I have just made between genre of an edited stream of consciousness and the genre of a piece of academic writing such as a PhD or a research paper, is a distinction between something which has merely instrumental value as evidence of something else (a thought process) and something which aims to have intrinsic value. Sketches, notebooks, and journals (or lab books) can also have the instrumental value.
So the first way of addressing problematic LLM use by students is to make very clear:
The kind of thing an LLM produces is not what we asked for. We don't want a polished answer but evidence of you thinking through the problem.
The second is to rethink our assessment instructions, and how we prepare students for those assessments, to make very clear what it is that we are looking for. We need to rewrite our assessment criteria in terms of the cognitive processes for which we were taking essays to be evidence. For example, we need to say that you demonstrate understanding of X not by simply saying what X is but by showing your own initial misunderstandings and how you corrected them.
I am reminded of a story that used to circulate in Oxford about David Wiggins. He was marking a very poor exam script and the student had crossed out the third essay and written 'This is rubbish'. Wiggins remarked: 'If he had done the same with the other two essays, I would have passed him.'
Perhaps our mistake is not asking for all the false starts, deletions, and corrections. I am not suggesting that this would be an 'AI proof' mode of assessment, but that it would clarify to students, and to the world, that what universities try to teach is something of a different order than the ability to produce text indistinguishable from an LLM output.