profile

Innovating with AI

How I think about AI safety today


Hey Reader,

It’s been a strange month for AI safety – starting with a round of alarm from current and former Anthropic employees about “recursive self-improvement” (more on that in a moment), and ending with…

Anthropic releasing a new frontier model, and the U.S. and China doing nothing in particular in the “AI race.”

Today I am going to share my thoughts on the more sensational side of this story, but more importantly, I’ll talk about how it does (and, largely, does not) affect daily life for professionals using AI (or helping their teams or organizations adopt AI safely).

The big takeaway for me from September is that AI doom is great for people who seek attention and “eyeballs.” That does not necessarily mean the claims are valid, accurate, or worth our limited time and attention.

Like pretty much everything in life, it’s largely up to you as an individual to stay calm, take it slow, and thoughtfully navigate a rapidly changing landscape. That includes understanding when maybe someone is pushing your emotional buttons in a not-totally-honest way.

At the same time, even if no new AI improvement ever takes place again, simply adopting the AI that exists right now is going to be a decade-long project that dramatically improves the efficiency and productivity of most organizations. Doing it in a human-centered, careful and thoughtful way is probably the most important thing you can focus on for the next few years. The process is going to be disruptive and at times emotionally taxing. It’s also generally going to make the things we do every day better, faster and cheaper.

So, my basic view – that AI will produce net positive effects in the form of better, faster and cheaper services and products for just about everyone – hasn’t changed much in the last three years.

The debate – around sci-fi doom fantasies (mostly silly), cybersecurity (real and important) and the environment (a mixed bag, depending how each country generates electricity) – is shaping things around the margins of that big idea. But it’s not really changing the big idea.

With that, let’s dive into this month’s most high-profile claims.

What to make of “AI doom” claims?

An Anthropic employee quit his job and his reasoning went viral. Another employee joined the conversation too. The key quotes:

“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
​
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

You can see why this would make a good viral news clip. Reminds me of the old saying in journalism – if it bleeds, it leads. 😑

The key here is that “self-improving superintelligence” is not real yet. But it is true that the big AI companies are pursuing the idea.

The gist is that you can use AI to build better AI faster. In some ways this is already happening. Every software developer uses AI to develop new software now. That extends to people who are building software for OpenAI, Anthropic and other big AI companies, like Meta and Google. So to some degree, AI can already speed up its own improvement.

The “next level” is setting up ways that this process can involve less human touch. The perceived risk is that you might get a situation where the AI model “runs away” and you, as the human creator, lose control of it to some degree.

I want to emphasize, though, that this is more like a theoretical sci-fi plot at this point. There is no evidence this has ever happened or is close to happening. It is not implausible, but it is also not real today. It’s sort of like a worst-case natural-disaster scenario like the “Big One” earthquake: possible but not necessarily going to happen on a timescale that matters to us.

So, it is a good idea to demand safety and regulatory procedures to be put in place as the AI companies pursue a more autonomous type of AI. It is also important to be aware of the actual reality today, which is…

The “10% chance” numbers are pure slop

Maybe it’s just because I’m a statistics nerd, but people saying there’s a “10% chance” of something bad happening – when there is no data or historical evidence to back that up at all, other than people’s opinions – rubs me the wrong way.

I see people saying “I believe there’s a 10% chance of AI killing us all” and I hear “I believe there’s a 10% chance of aliens visting this decade” or “I believe the world is going to end in 2012” or whatever other random conspiracy theory you want to use to fill in the blank.

In order to put a number on a prediction, you need at least two things:

  1. A clear definition of the event you are predicting
  2. A historical record of similar events that you can compare to the defined future event, so you can compare the circumstances and guess at a likelihood of this future event happening

You can do this with coin flips or poker easily. There is a clearly defined event and a nearly perfect historical record. The possibilities are finite. You can also do it with more complex systems like the weather or the outcome of an election, but you get much less precision because the systems are much more complex.

With the apocalypse predictions, you have none of the pieces necessary to express probability. So these guys who are running around saying there’s a 1-in-10 chance of every human dying need to be viewed with extreme skepticism. I understand that is a popular view among AI researchers. But they are irresponsible in framing it as anything but a “bad feeling” about the future.

(There is a line of reasoning that “since these people are smart we should trust them.” Fair to some degree, but “smart people” get things wrong all the time, and we still need to verify that their claims make sense.)

Bad feelings sometimes turn out to be right. They are sometimes totally off base. Regardless, they cannot be expressed as numerical claims, and you shouldn’t trust people who try to put numbers on vibes.

Since the claims are not based on any actual data, it is hard to figure out what to do

The AI companies have generally agreed that it is a good idea to give third-party observers access to their work. I think that is a great plan. However, as in many past AI news cycles, it seems like Anthropic is saying nice things, many people in the industry are agreeing verbally, and then nobody is changing anything. We should be skeptical of words that aren’t accompanied by action.

A reasonable analogy for third-party AI observers would be something like the Securities and Exchange Commission (watches public companies to make sure they don’t do fraud) or the Environmental Protection Agency (watches factories and power plants to make sure they don’t pollute too much). At the very least we get a public record of what’s going on inside the biggest American AI companies.

But I want to reiterate that:

  • This has been talked about but hasn’t happened yet
  • Unless a government mandates it, there’s little incentive for any AI company would grant significant public access to their internal workings
  • Observation and regulation don’t necessarily mean a disaster would be averted. Government agencies fail to prevent lots of bad behavior (not to mention accidents and epidemics).

So, industry insiders are talking about a pretty minimal amount of oversight, and even that isn’t a sure thing. At the same time, it’s not even clear what oversight would be needed to avoid the scenario that the Anthropic employees were imagining, because that scenario is not actually defined in a meaningful way. It’s a very vague nightmare.

The AI companies are not the only ones creating AI

This point is often brought up in the context of Chinese AI labs. Several Chinese AI companies are almost on par with Anthropic and OpenAI in terms of their most advanced models. (Specifically, Kimi K3 from China’s Moonshot AI currently ranks third on many charts, right behind the newest Claude and GPT models.)

So if the American AI companies pause for 6 months they are pretty much certain to “lose the lead” if Moonshot AI does not also pause. Maybe Moonshot will take the lead soon no matter what, but certainly it would be reasonable for someone who is a fan of democracy to want the American or European AI models to be in the lead.

I want to expand on this.

Even if zero for-profit companies were developing AI models, it would be essential for governments and militaries to continue to do so. Like the airplane, AI is both a consumer technology and a military technology. The airplane has been a huge net-positive for humanity; it can also be used to attack people in war. We should think of large language models in a very similar way.

And this is where the Anthropic researchers’ argument really breaks down. Let’s say, as a thought experiment, every company immediately agrees to pause AI development. It would still be irresponsible for major governments – the U.S. and China but also France, South Korea and many others who have the ability to build LLMs – to stop working on better AI. Even if the sole use of AI is cybersecurity, it’s essential – because the hackers already have really good AI models and have every incentive to keep trying to hack into banks and other critical institutions and infrastructure.

To bring back the airplane analogy – imagine an army with airplanes fighting one without. It wouldn’t be a close matchup.

So, even in the absence of for-profit AI development, there’s no plausible path to a worldwide AI development pause. You could actually argue it would be unsafe to pause because it would allow hackers to get the upper hand.

A since governments are effectively forced to continue AI development, the perceived risk (AI “running away” from our control) doesn’t go away regardless of what the current leading AI companies do.

What we can actually do today

Which brings me back to the thing I am actually focusing on in real life – and that I encourage you to focus on too:

Helping the people around you – at work and in daily life – adopt AI in a way that is responsible and sensible. Thinking about the humans first, and how making things better, faster and cheaper can benefit those around us. Yes, there will be disruption – but the nature and size of that disruption is largely within our control as people who work within organizations that are adopting AI.

Every single one of us can work hard to provide AI leadership in the organizations we care about. This does not require an international treaty. It is literally part of the stuff we wake up and do every day. And the more involved and caring we all are, the more likely each organization we touch is to do a good job handling this new era of tech.

Until next time,

Rob Howard
Founder, Innovating with AI


👋 Thanks for being a part of Innovating with AI. You’re in good company – some of our 175,000 readers work at Google, Apple, Microsoft and IBM… plus NASA, Morgan Stanley and the NBA.

❤️ Did you love (or hate) today’s newsletter? Reply to this email to tell me. I personally read every single message (without AI).

📬 Share this post (no paywall). If a friend forwarded you this email, subscribe here.

Innovating with AI

Coaching, community & curriculum to help everyone thrive in our AI‑powered future.

Share this page