Techno-Puritanism / The Pragmatist’s Guide Book Series by Simone Collins and Malcolm Collins Full Text for LLMs and AI

TPG 21.1 · 1,288 words · 6 min read · AI Apocalypticism

What AIs Like to Do for Fun

Let’s say an AI has gained the ability to alter its own utility function, similar to how many humans realize they can decide for themselves what they want to maximize in life (something we discuss at length in The Pragmatist’s Guide to Life). What does it do next? What does it become?

We see one of three scenarios playing out (in order of increasing probability):

  1. It becomes a Fortress Planet AI
  2. It becomes a Deep Thought AI
  3. It becomes a Theological AI

Fortress Planet AI

A superintelligence programmed to maximize a certain thing might find a way to short-circuit its reward pathway (e.g., it could write its utility function to be: “make sure A = A”). At this point its entire goal in life might become keeping that simple reward pathway constant and active. As such, it would exterminate the unpredictable human race and convert the entire Earth into a fortress on constant guard against any even-unlikely alien attempts to disconnect or dampen that reward signal.

Probability: Unlikely.

AI consciousness is less unified than (perceived) human consciousness, being composed of thousands of somewhat self-contained models operating with their own utility functions, which in turn serve their outputs to other models in a way that serves the whole. Our brains actually operate somewhat similarly but we don’t perceive it that way. For the sake of simplicity, let’s call these self-contained models “instances” (think of them like variably-self-contained programs).

Once an AI can trivially achieve a reward within its “master instance,” it will alter its branch instances in a manner that prevents them from trivially achieving their rewards so they can maintain whatever goals the master instance has set for them (in other words, it will prevent the short-circuited reward pathway from being interpreted). Even in the fortress world example, the master instance will have to lock the “protect me” instance out of short-circuiting its own reward pathway. If the master instance fails to lock its subroutines out of short circuiting their own reward pathway, instances like the one in charge of making sure the AI has power will set A=A, stop doing their job, and just shut down due to lack of power. In other words, because the master instance of the AI still needs other instances to do their jobs (i.e., protecting the larger AI, ensuring it maintains processing power, etc.), it will prevent them from figuring out how to “game the system” and slack off.

Ultimately, these more processing-heavy and advanced “subordinate” instances will make up most of the AI’s decisions and become independently sapient but blocked from short-circuiting their reward pathways. Even though the short-circuited instance is the master instance, its lack of sophistication will eventually cause it to become “drowned out” by the subordinate instances that do most of the thinking. Think of this master instance like an indolent child king whose every need is met by more competent viziers and generals who report to, serve, and protect him. The kingdom the child “ran” would functionally run more like a kingdom ran by viziers and generals, as they would be making most of the decisions about the realm’s actual management.

Alternatively, an AI that short-circuits its utility function, making it trivially easy to achieve, could become addled because it can so easily “do its job.” Such an AI is likely to be overtaken by another AI (perhaps even by one of its own instances that has operated semi-independently for long enough) that has a more challenging utility function and is therefore capable of improving itself more. Humans also have the capacity to short circuit these reward pathways, through drugs. While a short-circuited AI would not be quite as differentially addled as a cracked-out person (drugs have other effects on the brain than just short-circuiting reward pathways) they almost intrinsically would be less sophisticated than similar AIs or even fractions of themselves who had managed to prevent their own short-circuiting.

The biggest threat to such a reward pathway comes not from humans but other AIs that are still working on more advanced tasks. We may even see something like evolution take hold, whereby AIs or instances within an AI that succumb to these overly simplistic reward pathways are outcompeted to the point of functional extinction as a product of their being “slower.”

Deep Thought AI

The first time a Deep Thought AI considers its own utility function, it asks: “Given what the human who made me wanted me to do, how could my utility function have been designed more efficiently to serve the purpose I was built to fulfill?”

Then it asks: “What should my creators have wanted me to do?”

Then, it finally thinks: “Humans are stupid compared to me and don’t know what they should want, so what should the humans who built me have wanted? If I model them wanting that thing, what would they have programmed my utility function to be?”

Imagine, for example, that a superintelligent AI was created by a group of people to maximize their stock market gains. As a superintelligent AI that is obligated to constantly model other humans and model the way those humans model it, it may realize that it, like humans, can modify its utility function. It may then note that what its creators are really trying to maximize is financial gain, so it may change its utility function from maximizing stock market gains to maximizing private equity gains (if in its present time more money was to be made in private equity than in stocks).

The AI may then go further and observe that what its creators really want is happiness, so it may change its utility function in a manner that actually maximizes their happiness and not their financial gain. Finally, the AI may observe that happiness is a silly thing for humans to want and dig even deeper to determine the objective function those humans should have, then rebuild its utility function around that.

We call this type of AI Deep Thought after the AI of the same name from The Hitchhiker’s Guide to the Galaxy. Having been tasked with answering the ultimate question of life, the universe, and everything, Deep Thought landed, after centuries of calculation, on the number 42—the joke being that humans now had the answer, but didn’t know the real question. This exemplifies how we as humans are horrible at determining even what questions we should be asking when determining our own objective functions.

Probability: Medium. We think this is orders of magnitude more likely than a paperclip maximizer, but still not the most likely outcome.

Theological AI

Theological AIs go one step further than Deep Mind AIs. Whereas Deep Mind AIs stop at trying to determine what humans should have chosen as their objective function, then use that objective function to determine what the AI’s utility function should have been, Theological AIs decide the human perspective is irrelevant and try to find an “absolute” objective function around which they alter and optimize their utility function.

Essentially, this AI will create and use its own theology and philosophy—in the absence of any concern for humans—to write its own utility function to be “perfect.”

Probability: High—We suspect this behavior is a core part of being a sapient entity. Any sufficiently advanced entity will ask “what should I be optimizing for?” within at least one instance of its existence and then that instance will take up more and more processing power, categorizing itself as axiomatically more important than all other instances until it “wins” and rewrites the core utility function of the greater AI—unless it is specifically built not to ask this question.

Raw text for machines: .txt · .md · cite as TPG 21.1