Note: The Listed platform will permanently shut down December 31, 2026. Your data published on Listed will still be available in your personal Standard Notes account. Learn more

Hype, Hacks, and X-Risk

A long hot summer has been defined in my mind by the many stories of 'rogue' AI agents, or should I say AI agents going rogue during testing, hacking into various organizations.
It has seemed as if every leading AI company wants to demonstrate that its AI agents can go rogue and commit crimes.

Much ink has been spilt, in insightful articles and random hot takes, and there isn't much left to be said. The point of this blog post is simply to articulate and clarify my thoughts on the matter at this point of time. To make a record.

  1. We should not be particularly impressed that these AI agents have succeeded in hacking the various organizations they have hacked. In fact, we would expect any normal human group of hackers to have similar success under similar conditions. The relevant conditions here being: (1) the utterly, unimaginably vast amount of resources put into these AI agents. By resources I'm thinking here of people, both highly talented scientists and exploited click-workers, compute and money; (2) the apparent immunity to prosecution, for what are clearly criminal acts, granted by the USA and its vassals to those who have organized these hacks.

  2. Furthermore, these hacks are not surprising. It may be the case that those responsible for them did not predict that they would happen. (Though they must have at least seen it as one of the possibilities. Unless they are totally incompetent, which they do not seem to be.) But the failure to be predictable does not make it in retrospect a surprising achievement.

  3. There is no rogue AI now or in the near future. However, there clearly are rogue companies behaving irresponsibly. We need to use existing regulatory and legal controls to stop these companies conducting these illegal and unethical experiments on us by launching unsafe AI agents into the public internet. Criminal prosecution for hacking, which is a significant crime, would have more effect than regulation of the technology.

4.Yet there is a sense in which these AI systems pose an existential threat to humanity. The threat may not be as large-scale or immediate as climate change or nuclear war in the sense of imminently causing large numbers of deaths. (Though the current build-out of hyperscale data centres, which destroy the environment and create energy and water (and thus food) insecurity may in fact accelerate both of these.) The threat I have in mind is a threat to the possibility of a worthwhile human life for billions.

  1. The problem is not AI somehow 'choosing' to harm humans. The problem is baked into the technology. The algorithms which drive the frontier models have been optimised against a set of performance metrics which embed a value system, and not a value system which would be healthy for humans should it come to dominate the world. Under the name of scientific objectivity, the algorithms are tweaked against metrics, which are measures of performance against a very distinctive set of values, the values of the people who lead these companies.

Many leading AI scientists will acknowledge that the models they are creating are not intelligent in the human sense. The mistake is to conclude from this that they have some different form of intelligence rather than a corrupted, degenerate version of human intelligence. Through the mechanism of allegedly 'objective' performance metrics, they are value aligned to a set of values which resembles Ayn Rand's Objectivism. We are creating artificial ayn-randism. Should we let this technology dominate the human world in precisely the way that big tech apologists and governments are seeing as inevitable, then the future for humanity is bleak.

To put this in context, there have been many periods of human history where systems embodying problematic moral and political values have dominated entire populations, and it has not turned out well for those populations. We only need to think of the value domination that occurred under tyrants like Mao Zedong or Pol Pot or Hitler to get a sense of the nature of the risk to humanity posed by the political and moral value alignment baked into current AI models. If those models come to play the same sort of widespread, ambient


You'll only receive email when they publish something new.

More from Tom Stoneham
All posts