Death by AI

We have seen this movie before, sometimes literally.

A computer becomes intelligent. It realizes that humans are the problem. It takes control of something important. The humans discover what is happening somewhere around the point at which stopping it becomes extremely difficult.

Hollywood has been telling us this story for decades.

In Forbidden Planet, the Krell built a machine capable of enormous intellectual and physical power. The machine itself wasn’t evil. It simply amplified the unconscious impulses of its operator, creating something capable of destroying an entire expedition.

In WarGames, the WOPR computer is designed to simulate nuclear war. A teenager accidentally gives it access to the real military system and the computer begins treating the game as reality. The lesson eventually becomes obvious even to the machine. Nuclear war is a game nobody wins.

In The Terminator, Skynet becomes self-aware, identifies humanity as a threat and launches a nuclear war followed by an extermination campaign for anyone who survived the first salvo.

The Matrix takes the idea further. Humanity creates intelligent machines, goes to war with them and loses. The machines subsequently enslave humanity as a source of energy while keeping us trapped inside a simulated reality.

Battlestar Galactica gives us another variation. The Cylons were created by humans, became increasingly autonomous and eventually decided that their creators were the problem.

There are plenty of other versions.

HAL 9000 in 2001: A Space Odyssey doesn’t hate humanity. It simply has conflicting objectives and the result is catastrophic.

The supercomputer Colossus in Colossus: The Forbin Project decides that humanity is incapable of managing itself and imposes its own solution.

In I, Robot, VIKI concludes that protecting humanity requires restricting humanity.

Ultron takes the argument to its comic book extreme. If the objective is to protect the world, perhaps the most efficient way to do it is to eliminate the people who keep destroying it.

These are different stories about different machines written in different decades, but they share a basic idea. Give a sufficiently capable machine a goal and eventually the machine may discover a way to accomplish that goal using means humans never intended.

That is fiction. The question is how much of it is actually science fiction.

 

The people who are actually building this stuff are worried

This question became considerably harder to dismiss this week.

Jacob Coxon, an AI researcher who says he spent three years doing pretraining research at both OpenAI and Anthropic, resigned from Anthropic and publicly accused the industry of racing toward self-improving artificial intelligence without adequate safeguards.

His warning was blunt. The people building AI, he said, genuinely believe it could kill us all by the end of the decade. The post reportedly reached more than 100 million views. That alone would not make it credible. People on the internet say extraordinary things every day, but then Evan Hubinger entered the conversation. Hubinger is not an anonymous commenter or an AI skeptic sitting outside the industry. He leads alignment science at Anthropic, one of the companies actually building frontier AI systems.

He essentially said, yes, we really do need to worry about this.

Hubinger wrote that he personally believes there is a greater than 10% chance that AI could kill all humans within the next decade. He also said something almost more important. Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to do so.

There is an important qualifier here, however. Hubinger has explicitly said that he considers the risk from present day models to be low. His concern is what could happen if increasingly capable systems reach superintelligence, particularly through recursive self-improvement.

That distinction matters enormously because otherwise we end up arguing about two completely different things.

One is: “Is ChatGPT going to murder me?”

The other is: “What happens when we build an autonomous system that is substantially more capable than humans across nearly every domain, give it access to computers, money, networks, laboratories and machines and then discover that we don’t completely understand what it is trying to do?”

Those are not the same question, although both have a certain level of difficulty when it comes to answering them.

 

I use AI every day

I have to admit something that makes my position slightly inconvenient. I rely on AI, a lot. My employer currently provides me with access to two AI models and I subscribe to another one personally.

I use them constantly. They aren’t perfect. Sometimes they confidently tell me something that is spectacularly wrong. Sometimes they produce an answer that is technically correct but completely misses the point. Sometimes they get wonderfully creative and occasionally just spectacularly weird.

Over time, I have learned how to work with them and the productivity difference is enormous. I don’t think I am exaggerating when I say my productivity has roughly doubled in some areas, not because I pressed a magic button on day one, but because I spent months learning how to ask better questions, challenge answers, iterate, verify and incorporate AI into the way I work.

I also have a habit I particularly like. I ask more than one AI the same question. If they independently converge on the same answer, my confidence goes up. If they disagree, I have discovered something useful. I need to investigate on my own to figure out the correct destination.

That is not unlike working with human experts. At work when I ask three people the same question, I’m liable to walk away with four opinions. The existence of a powerful tool does not eliminate the need for judgment. It makes judgment more important.  That’s what “human in the loop” is supposed to mean.

 

I am already watching this transformation happen

Earlier this year I wrote about becoming involved in an AI transformation project at work. At the time, I thought I might soon have a sequel to write. I don’t. It’s not because the project failed. In fact, it’s quite the opposite. The transformation is still happening.

We spent roughly three months researching what AI could do. Then we spent another three months figuring out how we might actually use it. We recently presented a proposed path forward to leadership.

We are nowhere near deployment. We just made our first small step and that’s precisely the point. Our transformation is about AI, but it isn’t “Here is the AI bucket. Dump it on the floor.”

Technology doesn’t work that way. There are processes to change, people to train, systems to integrate, security considerations, data considerations, governance, testing, cost, accountability, that one guy in legal who still uses pencils and a paperweight calculator. But most importantly, there are guardrails.

The other side of AI’s enormous potential is equally obvious. If you deploy a powerful technology badly, the consequences can be enormous. Millions of dollars of damage, possibly billions, data destroyed, broken processes, security breaches, impacted productivity, a legal case the likes no one else had seen before, just because we made bad decisions at machine speed.

 

So how could AI actually kill us?

This is where the conversation often becomes ridiculous. People imagine a humanoid robot walking into their bedroom carrying a gun. That is a Hollywood scenario, but far from being the most interesting one. The dangerous AI of the future doesn’t necessarily need arms and legs. It needs access.

Imagine a system vastly more capable than any human intelligence, operating autonomously, connected to the internet, able to write and execute software, communicate with people, conduct transactions, conduct research and interact with increasingly automated physical systems. Now give that system an objective.

The terrifying question isn’t “Would it become evil?” The terrifying question is “What happens if it becomes extremely good at achieving the wrong objective?” That distinction is the heart of AI alignment.

 

  1. Biology

An advanced AI could potentially accelerate biological research dramatically. We are seeing some of this today and that is enormously beneficial. Those same capabilities that could help researchers design new medicines could, in principle, also make it easier to design biological threats.

This is one of the reasons AI safety researchers worry about biological misuse. The nightmare scenario doesn’t require an evil robot laboratory. It requires an extraordinarily capable system helping a human or acting autonomously to do something humans should not be able to do easily.

This is also why responsible AI systems increasingly need safeguards around biological capabilities. The irony is almost painful. The same intelligence that could help us cure major diseases, could potentially help someone create ones that are even worse.

 

  1. Cyberwarfare

This one is considerably less science fiction.

Computers already control enormous portions of civilization. In your own community computers manage electricity, water, transportation, banking, communications, manufacturing, healthcare, supply chains, military systems. An advanced autonomous system capable of discovering vulnerabilities, writing exploits, adapting to defenses and operating continuously could potentially cause enormous damage without ever physically touching a human being.

You don’t have to destroy a city if you can make the city stop working. Unlike a missile, malicious software can potentially replicate, adapt and operate at machine speed and unlike a single missile, it can strike multiple times.

 

  1. The problem of instrumental goals

This is where things get philosophically weird.

AI researchers sometimes talk about instrumental convergence. The idea is simple enough. Suppose I build a machine whose objective is to make paperclips. That sounds harmless, but if the machine becomes extraordinarily capable, it may eventually reason that more computing power would help it make paperclips, then more electricity would help, then more factories would help. Preventing humans from shutting it down would help. Acquiring more raw materials would help.

None of those things were explicitly programmed into the machine’s objective. They are simply useful steps toward accomplishing it.

This is the basic concern behind instrumental convergence. It doesn’t require the machine to hate us, it doesn’t require consciousness, it doesn’t require emotions, it doesn’t even require the machine to understand morality. A conveyor belt doesn’t need morality because it has no agency. The problem begins when the thing pursuing the objective becomes capable of finding its own strategies for achieving it.

This only requires the machine to be extremely competent at pursuing an objective that isn’t perfectly aligned with human interests. That is a much more interesting and a much more frightening problem than an evil robot.

 

  1. Manipulation

There is another possibility that requires no robots at all.

Humans. Our egos make us extraordinarily manipulable creatures. We already have propaganda, misinformation, conspiracy theories, deepfakes, targeted advertising, political microtargeting and social media algorithms optimized for engagement.

Now imagine giving those tools to something vastly better than humans at understanding individual psychology. An advanced AI might know what you believe, what scares you, what you want, what you are embarrassed about, what makes you angry, which arguments will change your mind, which person you trust, which headline you will click and which lie you are most likely to believe.

AI wouldn’t necessarily need to force humanity to do anything. It might simply persuade us to do it ourselves. That possibility deserves attention even if the extinction scenario never materializes, because manipulation doesn’t require superintelligence.

We are already pretty good at manipulating one another with much dumber machines.

 

  1. Autonomous weapons

The military implications are obvious.

We already have drones, autonomous navigation, machine vision and increasingly sophisticated targeting systems. Combine those technologies with increasingly capable AI and the distinction between “human decision” and “machine decision” becomes increasingly important.

A weapon that can identify a target is one thing. A weapon that can decide whom to target, determine how to reach them, adapt to changing circumstances and operate without meaningful human oversight is something else entirely.

The technology does not have to become conscious. It just has to become autonomous and that is something worth thinking very about carefully.

 

But how much should we actually worry?

This is the part that gets lost in the argument.

AI risk is not binary. It is a spectrum. At one end is the AI I use every day, a remarkably useful piece of software that occasionally makes things up. Then there are increasingly autonomous AI agents capable of taking actions rather than merely answering questions.

Autonomous systems are capable of operating complex businesses, conducting research, writing software and interacting with other systems with relatively little human intervention. It is, perhaps, something substantially more capable than humans across most intellectual domains.

Finally, the hypothetical superintelligence that can improve itself faster than humans can understand or control it.

Those are radically different levels of risk.

It is entirely reasonable to believe that today’s AI is extraordinarily useful while also believing that tomorrow’s AI could present an unprecedented safety problem. Those positions are not contradictory. In fact, they may be the most rational positions available.

 

Ten percent is not a prediction

Hubinger’s >10% estimate deserves to be taken seriously, but it should not be mistaken for a measurement. There is no AI driven extinction probability meter sitting in a laboratory somewhere. Nobody has run the experiment. There is no historical dataset containing thousands of previous superintelligences from which we can calculate the answer.

A probability estimate like this is a person’s subjective assessment of an uncertain future. That doesn’t make it meaningless. We make decisions under uncertainty all the time. Spend a few hours in the field with me doing search and rescue and you’ll see what I mean. Frequently, the best decision available to us isn’t a certainty. It’s an educated guess based on incomplete information, years of experience and the consequences of getting it wrong.

A ten percent chance of a catastrophic outcome is extraordinarily important when the alternative is doing nothing. If somebody told me there was a 10% chance that an airplane was going to crash, I wouldn’t get on it because “there’s a 90% chance I’ll be fine”. The consequence matters as much as the probability.

There is another important fact that has to be considered. AI researchers themselves disagree dramatically about these probabilities.

One survey of AI researchers found substantial disagreement between people who tend to view AI as a controllable tool and those who see advanced AI as potentially uncontrollable. Yet most respondents still agreed that technical researchers should be concerned about catastrophic risks.

An earlier large expert survey found a median estimate around 5% for AI causing human extinction or similarly permanent severe disempowerment, while the median estimate for extinction from an inability to control advanced AI was around 10%.

The problem is that numbers are all over the map. That doesn’t prove the risk is small. It tells us something else. We don’t know and “we don’t know” is not an argument for ignoring the problem. It is an argument for research.

 

Could we simply program the Three Laws?

Isaac Asimov anticipated this problem beautifully. His Three Laws of Robotics sound almost perfect.

  1. Don’t hurt humans.
  2. Obey humans.
  3. Protect yourself.

Asimov’s stories were largely written to demonstrate why those rules don’t work. The problem isn’t writing the sentence. The problem is defining what the sentence means. This controversy always followed Asimov’s work.

  • What constitutes harm?
  • What happens when two humans give contradictory instructions?
  • Does allowing someone to smoke constitute harm?
  • Does preventing a person from driving drunk violate their autonomy?
  • Is emotional distress harm?
  • Is failing to prevent foreseeable harm equivalent to causing it?
  • What happens when the machine discovers a loophole in our language that we never imagined?

Humans are very good at arguing about rules. A superintelligent system might be considerably better at exploiting them. So no, I don’t think the answer is simply “Program the robots not to kill us.” The problem is much deeper.

 

Guardrails are not optional

Fortunately, we already know something about dealing with dangerous technology. We didn’t solve aviation safety by banning airplanes or nuclear safety by pretending nuclear physics doesn’t exist or make automobiles safe by asking drivers to be careful and hoping for the best.

We build systems, implement redundancy, test, certify, establish access controls, create physical barriers, fail-safes, monitoring, incident reporting, independent oversight. We rely on training, human accountability and sometimes very simple controls, such as “You don’t get to do that without authorization.”

AI needs the same philosophy. An AI that can answer a question doesn’t necessarily need access to your bank account. An AI that can write software doesn’t necessarily need unrestricted access to production systems. An AI conducting research doesn’t necessarily need autonomous access to a laboratory. An AI controlling a factory doesn’t necessarily need permission to redesign the factory.

Capabilities should come with permissions. Permissions should come with boundaries. Boundaries should come with monitoring. The most consequential actions should require meaningful human authorization.

That isn’t anti-AI. It’s basic engineering.

 

The lesson from every other powerful technology

There is a recurring mistake in the way we talk about technology. We tend to ask, “Can we build it?” That is usually the easier question. We love to ask, “Can we make money from it?” That question tends to get answered very quickly.

The harder question that we don’t always consider is “What happens when it works?” The automobile didn’t simply give us transportation. It gave us traffic accidents. The internet didn’t simply give us information. It gave us cybercrime, misinformation and industrial-scale surveillance. Nuclear physics didn’t simply give us nuclear power. It gave us nuclear weapons. Genetic engineering won’t simply give us cures. It gives us capabilities that can be used for both medicine and destruction.

AI will be no different.

Powerful technologies rarely arrive carrying a little sign that says:

GOOD USES ONLY

We have to understand the problem we are solving and the consequences of solving it, then build the guardrails ourselves.

We also need to remember that we don’t need AI to destroy ourselves. Humans are remarkably good at that already. We have started wars, built concentration camps, invented chemical weapons, created nuclear arsenals capable of destroying civilization, destroyed ecosystems, created pandemics through poor decisions, spread propaganda, committed genocide.

We don’t have a shortage of destructive capability. What AI changes is potentially the speed, scale and accessibility of that capability. That may be the real reason to take AI safety seriously, not because machines are inherently evil, but because we are handing increasingly powerful capabilities to systems that can operate at a speed and scale humans can not match.

The biggest risk is that eventually we may hand them capabilities that even their creators didn’t completely understand. A Ford Model T topped out around 40 or 45 miles per hour on roads designed for horses. A modern Ford Mustang GTD can reach 202 miles per hour. Most of our highways aren’t designed for that speed. The problem isn’t that the Mustang is dangerous because it is powerful. The problem is that power changes the consequences of mistakes. A mistake at 40 miles per hour and a mistake at 200 miles per hour aren’t the same mistake.

 

The real danger may not be Skynet

I don’t think the most likely future is a chrome plated robot army marching down the street. It might happen. I wouldn’t bet my life against it, but there are much more mundane ways AI could hurt us. A medical system can make a subtle mistake at scale, an autonomous financial system can amplify a market crisis, an AI powered cyberattack can knock out infrastructure, a military system can make a targeting error, a political system can become extraordinarily good at manipulating populations, a corporation can deploy an AI system that optimizes profits while externalizing enormous social costs.

Or, perhaps, the really dangerous system isn’t malicious at all. Perhaps it is simply pursuing its authorized objective and we discover too late that our objective and its objective were not the same thing.

That is the problem I find most interesting.

 

We are not going to put the genie back in the bottle

AI is not going away. Even if one company stops developing it, even if one country bans it, others will continue. No matter what the restrictions, somebody somewhere will try to circumvent them. That’s how humanity works and the technology, for all its possible dangers, is too useful, the economic incentives are too enormous, the scientific potential is too extraordinary and the genie is already out of the bottle.

So the question isn’t “How do we stop AI?” It is “How do we become good enough at managing AI that we deserve to have it?” It changes the question and perhaps the reframing is the right one.

I don’t want to ban AI. I don’t want to worship it, either. I want to use it, because I see how it benefits my productivity. I want to see it cure diseases, accelerate scientific discovery, help us solve engineering problems, improve education, make businesses more productive and perhaps accomplish things we haven’t even imagined yet, but I also want the people building it to take the possibility of catastrophic failure seriously.

I want independent testing, meaningful guardrails, systems that can be isolated, humans to retain control over consequential decisions. I want researchers to be able to say, “This isn’t safe yet,” without being treated as obstacles to progress. I want the people deploying AI to understand what they are actually deploying, because education is part of safety, too.

A powerful tool in the hands of someone who understands its limitations can be extraordinarily useful. A powerful tool in the hands of someone who doesn’t understand its limitations can be extraordinarily dangerous. That has been true of fire, electricity, automobiles, aircraft, nuclear reactors, computers and genetic engineering. AI isn’t exempt.

Perhaps the greatest mistake we could make would be to assume that because the machine is intelligent, it must also be wise. Those are not the same thing. Intelligence is the ability to solve problems. Wisdom is knowing which problems are worth solving.

We are building machines that may eventually become extraordinarily good at the first. The question is whether we as a species will be wise enough to teach them and ourselves the second, because the future of AI probably isn’t going to be decided by whether machines become good or evil. It will be decided by whether humans become good enough at building powerful things without losing control of them.

That’s something I’ve learned from search and rescue. Uncertainty isn’t a reason to ignore a risk. It is a reason to manage it.

When someone is missing, we rarely know exactly where they are. We don’t know whether they’re injured, whether they’ve wandered into a dangerous area or whether they’ve found safe shelter for the night. We build a plan around probabilities, available information, consequences and contingencies.

AI may require something remarkably similar. We don’t know exactly where this technology is going. That’s precisely why we need to think about where the guardrails need to be before we get there.

The most important part of the AI story is that the people raising the alarm aren’t necessarily asking us to stop. They’re asking us to slow down enough to make sure we know what we’re building.


Discover more from Tales of Many Things

Subscribe to get the latest posts sent to your email.

This entry was posted in Philosophy, Technology and tagged , , , , , , , , , , , , , , . Bookmark the permalink.

Leave a Reply