The 'Problem' of AI Cheating

Some six weeks ago, Prof Andy Hamilton, a philosopher at Durham University who I have met a few times, published a short piece in The Spectator entitled 'Why isn't Durham University taking AI cheating seriously?'.

Hamilton shows great disrespect for students, misdiagnoses the problem, and offers a solution which addresses a different issue, one which pre-dates ChatGPT by at least 20 years.

Cheating Students

At one point Hamilton remarks:

Students have always cheated, but now they know they can do so with impunity.

Well, yes, a tiny minority of students have always tried to cheat. However, in my >30 years of experience almost all students do not want to cheat. They want to get good marks for the work they do. They want to be proud of their achievements and know that good marks achieved by cheating engender no pride. As Hamilton will know from his long experience investigating academic misconduct, primarily plagiarism, pre-ChatGPT, most students caught plagiarising are distraught by the idea that what they did was cheating: they made a mistake, took a shortcut, or didn't understand the principles of academic integrity. Since most universities operate a principle of strict liability for plagiarism, this lack of intention is no defence, but they still want us to know it.

Hamilton seems to to think like Glaucon in Plato's Republic, reflecting on Gyges' ring which makes the wearer invisible:

And this we may truly affirm to be a great proof that a man is just, not willingly or because he thinks that justice is any good to him individually, but of necessity, for wherever any one thinks that he can safely be unjust, there he is unjust. ... If you could imagine any one obtaining this power of becoming invisible, and never doing any wrong or touching what was another's, he would be thought by the lookers-on to be a most wretched idiot (trans. Jowett)

As a time-served academic philosopher, Hamilton knows that this is at best a minority view, at worst a travesty of the concept of reason. It also appears to be empirically false. When we observe our fellow humans, students included, we see that no one is perfect. As Hume put it:

Were one to go round the world with an intention of giving a good supper to the righteous and a sound drubbing to the wicked, he would frequently be embarrassed in his choice, and would find, that the merits and the demerits of most men and women scarcely amount to the value of either. ('Of the Immortality of the Soul')

Yet we also observe that almost everyone, almost all the time, prefers to do the right thing or follow the rules. Threats of punishment work at the margins, but for the overwhelming bulk of cases, people avoid what they know to be wrong. And the same is true with students and cheating.

Hamilton thinks the problem is that AI makes cheating easy (if you can pay for the best models, anyway). That is Glaucon talking. However, he also seems upset that AI doesn't merely make cheating hard to detect, but also produces work of higher quality than most non-cheating students can achieve:

... the lazy student cheats with professional-grade versions of AI chatbots such as ChatGPT or Claude, and gets a first-class result. The hard-working student uses no bots, or a non-professional bot honestly, thinking and writing for themselves, but gets a 2:1. That’s not fair.

It is interesting that Hamilton finds the cheating student getting a first more unfair than the cheating student getting a 2:1, even though standard graduate job offers only require a 2:1. In fact, having a first makes little or no practical difference unless one wants to pursue an academic vocation, though it might make one proud of one's achievement (at least, if one hadn't cheated to get it).

Diagnosing the problem

Two weeks before Hamilton decided to 'blow the whistle', an important piece of research on the subject was published in Science using survey data from 95,513 students across 20 'major public research intensive universities in the United States'. Using a technique called 'list randomization' which is designed to get data on sensitive statements in a manner which allows anonymity, the researchers were able to get estimates of 'AI cheating' by discipline. This showed an average across all disciplines of 9%, with a slightly lower average for STEM disciplines and higher for non-STEM. Humanities, which includes Philosophy, comes in at around 7%. Given the tone of Hamilton's article, and the general moral panic about AI cheating in universities, these are strikingly low figures.

However, the really important point about this research is the exact phrasing of the statement which was taken to provide evidence of 'AI cheating':

I have submitted AI-generated content as my own work in class, knowing it may not be allowed.

Words matter and the choice of 'may' is important. It means that what is being measured here is not just students who know they are cheating, not even those who believe they are cheating, but also those who do not know they are not cheating. The uncertain.

This means that the 9% on average who are being reported as engaging in 'AI cheating' should actually be broken into three distinct sub-groups by intention:

  1. Those who intend to cheat and use AI to do it
  2. Those who do not intend to cheat, but under pressure of deadlines use AI as a shortcut, knowing this isn't really doing the work 'properly'
  3. Those who do not intend to cheat, but are confused by mixed messages and unclarity about what is and is not acceptable use of AI in their studies

As should be clear, I think the first group is very small. Perhaps it is larger in Durham, but that would be a different issue. What is important is how we help - not punish - the second and third groups.

Take the second group, the 'shortcut' students. These students are the ones who may also not do the reading they have been set but instead try to find online sources which tell them what it is about. They know that what they are meant to do is read the text we have given, but that is hard work and they have many other demands on their time, from paid employment or sports to other academic assignments. So they get an AI summary and use that, lightly edited, in their essay. Before AI they would have sought other shortcuts - the internet is full of 'helpful' websites explaining the content of academic courses.

Universities can do something about this. Of course, they can't reverse the abolition of adequate maintenance grants, or the rising costs of food and utilities, or price-gouging of private landlords - that is for the government. But they can rethink the assumptions around assessment which generate the pressure of deadlines. For example, in my Department, we have a single deadline for all modules in a semester (by cohort). This makes it much easier for student services staff to manage and respond to missing submissions, extension requests etc. It also makes it easier for markers to manage their time. But those efficiencies come at a cost and that cost is borne by the students. Having to submit three (or more) essays at exactly the same time on the same day is considerably more stressful than having one a week for three weeks.

In fact, the old exam system always spread the exams out over a couple of weeks. The real negative impact of moving to 100% coursework is the stress we chose to cause students to make things more convenient for ourselves. Having extensions and self-certifications and late penalties mitigates the problem. Yet it is a problem of our own making and entirely for our convenience. If we went back not to exams themselves but to the staggered temporal structure of assessments that was a by-product of the exam system, we would do a lot of reduce the incidence of the second group of 'AI cheats'.

What then of the third group? This is where we really can trace the blame to the very top. Vice-Chancellors and their Executive Boards have fallen for the AI hype, partly due to pressure from governments desperately hoping that an AI future will solve the problem of achieving endless economic growth, partly due to quite overt lobbying by Big Tech, and partly due to a certain non-critical naivete in their own experience of AI. As a result, university bureaucracies the world over are sending out the message to staff and students that being able to use AI is a key employability skill and that we must make sure our students acquire this skill by 'innovation' in the curriculum. What students hear is that they should be using AI in their studies because that is the future of work.

Simultaneously, the same bureaucracy is saying that we must uphold academic standards and not use AI in ways that go against the principles of academic integrity. Which is good and true. Unfortunately, instead of stating clearly which uses of AI are academically problematic, under the veneer of 'academic freedom', they try to push that down to the level of the module leader to decide. The result, of course, is chaos and unpredictability for students. What is encouraged on one module may be banned on another, while a third doesn't say anything specific at all.

Effectively, we have created a situation where students are told that they should use AI (except where they shouldn't) and they are left to navigate that decision themselves. This is irresponsible in so many ways, but mainly because it is unfair on our students and erodes academic standards at the same time. It is the responsibility of the most senior and most visible of university leaders to sort out the mess and to do so very quickly.

The only way to do this is to make universities a safe space from the AI hype. Perhaps our students will need to be able to use AI in their future workplace, but that doesn't make it our responsibility to teach that, any more than we teach them to use email or a word-processor. There is a cogent and clear message that could be given, but it would have to be given from the top:

The point of a university is to teach cognitive skills, not shortcuts. If the Learning Outcomes of your module do not mention using AI, then you MUST NOT use AI in that module assessments.

If you are still worried about the first group, the intentional cheats, then I have a suggestion: randomly viva 5% (or 10% if you are really worried). This will have a deterrent effect. Of course, the AI cheat who prepares for the viva by actually learning enough to defend their AI generated essay will pass the viva. But then again, they will have done that by actually doing the work they tried to avoid, so not much harm has been done.

The value of anonymity

Hamilton's quick fix for dealing with 'AI cheating' is to drop anonymous marking. Presumably this is intended to address the unfairness of a 'second class mind' getting a first-class degree as a result of using AI. It presupposes that the marker who has taught the module will know which students are capable of writing a first-class essay, and which are not.

The presupposition is, of course, entirely false. Most university modules have at least 50 students on them and some several hundred. Lecturers struggle to learn their names. And how to assess their abilities, their potential? Through class contributions? But only if they turn up regularly (and the most able may be most bored by 'mixed-ability' teaching) or submit formative work (not AI generated!).

Setting that aside for now, let us consider the reasons for anonymous marking. Hamilton remembers the introduction of anonymous marking 'which happened because research showed that female students were discriminated against'. Yes, the operation of implicit bias is one reason to have anonymous marking, and it is not just female students who suffer. However, there is also a more straightforwardly academic reason: we mark the work produced not the person who produced it.

Of course there is a correlation between the abilities of the individual and the quality of the work they produce, but the most able sometimes write a duff essay and the average student sometimes writes an outstanding essay. Anonymous marking ensures that they get the mark the essay deserves. If it were not anonymous, the able student writing a poor essay would get too much benefit of the doubt, and the average student would be suspected of cheating when really they had just excelled themselves. Both are unjust outcomes.

What is problematic about this system of anonymous marking is that a principle of aggregation of the marks awarded to multiple assessments produced by the same individual is translated into a 'Degree' which is a form of marking or grading the person. That fundamental incoherence in university education is nothing to do with AI and far too big a topic for today.


You'll only receive email when they publish something new.

More from Tom Stoneham
All posts