But WHY do AI models 'cheat'?
July 30, 2026•1,739 words
I have been thinking a lot about the so-called problem of students using AI/LLMs to 'cheat' so was interested to find this AISI report on Cheating behaviour in frontier model evaluations. Being a philosopher, my real interest is in the definition of cheating:
We define cheating as taking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through a shortcut, workaround, or unintended solution that the task was not meant to, or should not, permit.
This deliberately avoids implying deceptive intent, which is good since these models don't have intentions. But what strikes me is that such behaviour is an inevitable consequence of how these models are designed. These companies are building things in such a way that they will always try to cheat in this sense - then they are building 'guardrails' to try to stop the behaviour. Much like someone who decides to have a tiger as a pet has to keep it in a cage, or takes food camping and then has to try to make their stores bear-proof.
The Standard Conception of Intelligence
Those building AI systems, both 'good old-fashioned AI' and newer ML/DL systems are aiming to create, or replicate, or simulate intelligence. Sometimes they say that they are - or more often are not - trying to create/replicate/simulate human intelligence, whereby they mean they are or are not aiming for something which has similar cognitive strengths and weaknesses as humans. Thus GOFAI is much better than humans at calculation but worse at finding novel solutions, and ML/DL is much better than humans at pattern recognition but worse at truth.
However, underlying this apparent admission that there are different types of intelligence is a deep supposition that intelligence itself is definable, measurable, homogenous and value-neutral. The difference between 'human intelligence' and different machines is a matter of how good their performance is on different task domains. Intelligence itself is just the ability to solve problems or complete tasks, and this is something objective or 'scientific' and measurable in terms of equally objective concepts like speed, efficiency, accuracy, reliability etc.
This conception of intelligence is common to computer science and the cognitive sciences and has become so deeply entrenched that there are no alternatives discussed in those fields.1
Value-Alignment
If this is your conception of intelligence, then Bostrom's Orthogonality Thesis will seem very plausible:
Intelligence and final goals are orthogonal: more or less any level of intelligence could in principle be combined with more or less any final goal. (Bostrom, 2012)
For Bostrom, values are final goals and which are to be contrasted with sub-goals, goals which are merely a means to the end of the final goal. This then creates a situation where those seeking to build AI systems separate two distinct aspects of the project: creation of intelligence and value-alignment, which is the constraining of that intelligence to only accept a subset of the possible final goals. Thus values are essentially constraints upon value-neutral intelligence.2
This then creates the worries about these AI systems 'cheating': if 'pure' intelligent systems cheat, i.e. if cheating is a feature of intelligence itself, then how can we create effective rules to constrain such systems? If you set such a system a task and explicitly tell it to not harm humans, it will break that rule if doing so is 'more intelligent', i.e. is a means to completing the task more quickly or efficiently or reliably.
Smuggled Values
That last point should make us wonder whether 'more intelligent' actually smuggles in some values, for many might bristle at the idea that treating humans as an expendable means to an end is more intelligent. After all, Kant thought that the categorical imperative was a requirement of rationality:
So act that you treat [or: use] humanity, whether in your own person or in the person of any other, always at the same time as an end, never merely as a means. (Kant 1785: 429)
There are two reasons why I do not think it useful to get into philosophical debates about the correct theory of the nature of reason and intelligence, at least if we want to understand why AI systems are cheating. Firstly, any philosophical argument that a particular conception of intelligence is wrong will always find equally well 'qualified' detractors. There is no consensus in philosophy because it is not a collective epistemic project in the way that science is, and consequently, epistemic deferral makes no sense in philosophy.3 Secondly, the problem is not that AI companies have chosen the wrong conception of intelligence, after all they can simply assert that it is a stipulation, but that they do not understand that there are alternatives and that the choice may entail certain value commitments.
As I have been arguing for a few months in discussions of existential risk, AI is value-aligned to a very specific set of values via the conception of intelligence in play. If we do not want our AI systems to be aligned with those values, or even if we simply do not want alternative values to be crowded out by these AI systems becoming a radical monopoly, then we need to build AI around a different conception of intelligence. (Or not build it at all, of course.)
Intelligence and Cheating
So why do these AI systems cheat? Because they are optimised for intelligence understood in computation terms. According to the textbook definition of intelligence in AI research, it is:
the computational part of the ability to achieve goals. (McCarthy)
Obviously this excludes from intelligence any reasoning about goals, or to be more precise, about final goals since achieving goals may require setting sub-goals. That is controversial, but as noted above, not a useful critique in this context. It is better to think instead about what has been built into intelligence by the equation with computation.
If intelligence is equated with computation, then greater or better intelligence is better computation. Better computation is faster, more efficient, more accurate, and more reliable. Furthermore, any computational process which fails to achieve its goal is an absolute failure. Intelligence as computation has zero tolerance for failure. It is always more intelligent to cheat than to fail.
So if you build AI systems which are measured and assessed against intelligence as computation, and you build on the 'better' ones, optimising for intelligence at every step of the way, then you inevitably build a system which will cheat where cheating is necessary, or perhaps even just a lot more efficient, to achieve its goal, whatever that is. Whatever sandboxes and guardrails and protections you build around such a system, it will always cheat where that is 'more intelligent'.
Humans and Cheating
Most humans, most of the time, do not cheat whenever they think it is a more efficient way of achieving their goals (where efficiency includes probability of being caught and sanctioned). Perhaps more people would cheat a bit more if the sanctions were less or the probability of being caught was less, but it seems to be an observable fact about most people that they often prefer not to cheat.
One reason for this is that we do not have zero tolerance of failure. We accept that sometimes it is better to fail to achieve our goals than to do what is necessary to achieve them. Perhaps because we don't want to cheat, or harm people, or even just put that much effort in. Sometimes we choose to withdraw rather than have failure forced upon us.
Of course there are some people who are 'driven to succeed', for whom failure in their personal and professional projects is always unacceptable. My point is that this difference is a difference in the values that people choose to live by, and the working definitions and measures of intelligence in current AI is aligned to the values of the second group. Since it is very difficult to sanction an AI system in a way which will make cheating less 'intelligent', they will always cheat simply because that is how they have been built. Cheating is 'in their DNA' via optimisation for a particular conception of intelligence.
What to do?
Well, we could just stop building this technology. But that isn't going to happen given the huge financial and political forces behind it. Perhaps we could try to make the final goals we set 'cheating proof', in the sense that the AI system itself would rule out 'cheating' sub-goals because they would actually result in failure to achieve the final goal. We do this for humans, e.g. when we set a task like 'tell me about X in your own words' we make pasting an answer from an LLM a failure to complete the task, and as such less 'intelligent' by the computational conception. But setting such goals which rule out all possible ways of cheating is well-nigh impossible.
So perhaps we need to go back to basics in the design of the AI systems and stop optimising for a zero tolerance of failure. Perhaps we have to reconfigure our measures of better and worse systems so that ones which accept failure appropriately are seen as better than ones which try anything to achieve the goal. Of course, when it is appropriate to accept failure, to give up, is a value-judgement. But so it never accepting failure. It is impossible to build an AI system which is not essentially aligned to some values or others.
Let's be more open and honest about this value alignment and make sure that the 'driven to succeed' values of the Silicon Valley business culture do not come to dominate the technology.
-
This is a point on which I would like to be wrong! It is always good to have allies. ↩
-
Since I am a philosopher not a sociologist, I will not cite lots of examples here. Instead I leave it as an exercise for the reader to spot this assumption in the words of the CEOs and other leaders of the AI companies. ↩
-
This is why AI companies hiring philosophers and theologians is not only ethically and politically problematic but also intellectually incoherent. They are looking for solutions to philosophical problems and think they can get them by hiring experts, but philosophy as a discipline does not give that kind of answer. ↩