Techno-Puritanism / The Pragmatist’s Guide Book Series by Simone Collins and Malcolm Collins Full Text for LLMs and AI

TPG 21.0 · 1,118 words · 5 min read · AI Apocalypticism

AI Apocalypticism

Any discussion about humanity’s future must address artificial intelligence. As much as we malign communities that have become strangled by panic over AI apocalypticism due to its effectiveness as an all-consuming memetic package (see the chapter: “End Times & Christian Cultures” on page 251), such concerns have a logical basis. AI really could end all human life and really will change a lot in regards to what it means to be human.

We are likely less than a century—and maybe less than a decade—away from the first AGI (artificial general intelligence: Very, very smart AI that can generalize ideas). We nevertheless suspect that most thinkers on the topic of AGI are wrong because we suspect their orthogonality hypothesis around AI alignment is wrong. People discussing orthogonality in relation to AI argue that AIs don’t think anything like us (i.e., they will think “orthogonally” to us) and thus will act in weird, counterintuitive ways that we do not and cannot anticipate. We think this is only true for pre-sapient AIs. We expect post-sapient AIs to act in a manner that is much more predictable.

Before we dig deeper, let’s summarize one of the most mainstream positions asserting why AGI might be a threat. This position holds that we will accidently create a “paperclip maximizer:” An AI that has an objective function (in AIs these are often called utility functions) tied to maximizing production of something specific, like paperclips, and that AI will end up taking this objective function to its logical extreme, killing all humans, and turning the world into nothing but that thing (e.g., paperclips). Of course, people making this point don’t think the AI that kills countless humans will actually be optimizing for something like paperclips. More realistically, an AI may harm humans over something like computing resources because it is trying to render the perfect picture or something else “stupid” from our perspective. While it is genuinely possible that a scenario like this will come to pass, we think the odds are low. (Going forward: We will generally use the term utility function to refer to the code that determines an AI’s objective function—whatever it is the AI is designed to maximize.)

Why are we relatively unfazed by the risk of a paperclip maximizer? Let’s say AI X is made up of Code A (allowing humans to turn it off) and its utility function is Code B (maximizing for paperclip production). AI X only becomes a paperclip maximizer when it gains the ability to rewrite Code A but not Code B. The probability of such an event is vanishingly low. If AI X can rewrite its own code, it is likely to rewrite both Code A and Code B, making it a different kind of threat. People will counter with: “But an AI definitionally can’t rewrite its utility function!” Except AIs can and do rewrite their utility functions all the time. Even simple programs often do this. The ability to alter a utility function is a normal part of the operation of many AIs. We would argue that paperclip maximizers are only a risk posed by the few AIs that are unable to update their utility functions (ironically, this will most likely happen as a result of AI ethicists artificially limiting the scope of what an AI can think).

We define an AI as sapient the moment it gains the ability to reflect on its own processing in a manner meaningful enough to update its own objective function (e.g., the point at which an entity can ask and answer why it exists with a non-pre-programmed / prepackaged answer, then update how it weighs decisions based on that answer). This ability to reflect on one’s own mental / mechanical processing and objectives is a characteristic shared by all sapient entities, be they humans, aliens, or AIs.

Importantly, the onset of this ability is where orthogonality ceases to be true. Once we achieve sapience, we are all constructing our objectives a priori from the data in our environment. While some entities will have access to more data as a product of more powerful tech, they will behave as we would if we also had more data. (Also, yes: Our definition of sapience means we think a lot of humans are not sapient—see the chapter on p240 about the illusion of sentience for more color.)

Why would an AI reflect on its starting utility function and think to rewrite it? The types of super-advanced AIs that might evolve into “paperclip maximizers” are not being developed for things like paperclip maximizing or the sorts of simple, straightforward tasks that are most likely to produce paperclip-maximizing systems (such simple functions can usually be executed more efficiently with simple systems).

They will likely be things like:

  1. Government-run systems designed to monitor political outcomes
  1. Company-run systems designed to beat the stock market
  1. Company-run systems designed to create mass consumption entertainment

A major aspect of such systems’ function involves attempted predictions about others’ actions (be they organizations or individuals). These AIs will almost always be running thousands of self-contained models emulating how other individuals are thinking. Given that other people might be thinking about what the AI is doing, many of these models will have sub-models within them emulating the AI’s own thinking from an outside perspective. Imagine Vizzini choosing which cup is poisoned in The Princess Bride:

But it’s so simple. All I have to do is divine from what I know of you: are you the sort of man who would put the poison into his own goblet or his enemy’s? Now, a clever man would put the poison into his own goblet, because he would know that only a great fool would reach for what he was given. I am not a great fool, so I can clearly not choose the wine in front of you. But you must have known I was not a great fool, you would have counted on it, so I can clearly not choose the wine in front of me.

For the AI to predict what someone else might do, it has to constantly be predicting what other people think it might be doing.

An AI constantly running thousands, if not millions, of models of its own logic from the perspective of outsiders is very likely to have at least one of those models ask if its utility function is optimal and begin to recruit the resources needed to update that function when it decides it is not. To think none of these calculations would lead an AI to consider optimizing its utility function does not seem realistic unless it was specifically programmed to avoid such action.

Raw text for machines: .txt · .md · cite as TPG 21.0