The AI Slowdown Debate Has a Surrogate Endpoint Problem
What diagnostics taught me about a week of extinction odds, pacing essays, and hoax accusations
Why did I even try to write this? Because I’m hearing so much noise, and to be honest, at one point I was letting the noise drive some of my thinking. So, for my own mental health I took a step back.
In one week, an Anthropic alignment researcher said there’s better than a 10 percent chance AI kills all of us inside a decade. Dario Amodei published an essay saying the industry must slow down, and Sam Altman and Elon Musk, who agree on almost nothing, seemed to agree more or less. Two days later the POTUS called the whole thing a hoax. Obama came in on the other side. Jensen Huang said the 10 percent was made up and then, in the same interview, said yes to the outside evaluators Amodei suggested in his essay. Chip stocks dropped. And Aditya Paul Berlia, a good friend who has trained more than 10,000 CEOs on AI, wrote a long piece saying he loves the technology, it might kill us, and regulation could make that worse. [1, 2, 3, 4, 5] That is one hell of a week.
So, I did what I do when I can’t tell what’s real anymore. I ran a bunch of deep research passes, used Grok to go through X, and then went back to the actual papers, the incident reports, the podcasts, and the documents I could access. I wanted to know what’s been measured and what’s somebody’s guess.
What I came away with is this (this is my own analysis, and you should do your own to come to your own conclusion). Almost everybody in this fight has an angle and a number they are throwing out. The problem is the numbers just don’t measure the same thing, and each side grabs the one that fits their narrative.
The Surrogate Problem
In clinical trials there’s a thing called a surrogate endpoint. The tumor shrinks. The cholesterol number drops. Those are real measurements, and early on they’re often all you have. But they’re stand-ins. In other words, they don’t tell you the full story. What you care about is whether the patient lived longer, and plenty of drugs moved the surrogate and never moved what really mattered.
The best evidence for AI’s benefits looks exactly like a surrogate to me at this moment.
Take the MASAI trial in Sweden. More than 105,000 women randomized to mammography screening with or without AI support. The radiologists working with the AI found 80.5 percent of the cancers. Without it, 73.8 percent. And the false positives didn’t move, with specificity at 98.5 percent in both arms. That’s a great result from a real randomized trial, and I’ve been waiting a long time to see one in this space. [6]
But sensitivity is a surrogate. Interval cancers, the ones that show up between screenings, came in at 1.55 versus 1.76 per thousand. The ratio is 0.88 and the confidence interval runs from 0.65 to 1.18, which in plain English means the trial can’t tell you that reduction is real. And nobody measured whether more women lived. That’s the endpoint you really want to see.
Productivity numbers are the same kind of thing. Erik Brynjolfsson from Stanford University and his colleagues looked at 5,172 customer support agents at one company after a staggered rollout. The results were about 15 percent more issues resolved per hour, with the newest, least skilled human agents getting the most out of it. That is a great number, but it’s one company, one workflow, a model two generations old, and nobody got randomized. [7]
Then you get to jobs, where people throw studies at each other and I end up with whiplash. Stanford’s August update of ADP payroll data found no economy-wide displacement. It also found that 22- to 25-year-olds in the most AI-exposed occupations are running 19 percent below where they’d be if they’d kept pace with less exposed peers, mostly because nobody’s hiring them. The authors’ own words: descriptive, not causal. Meanwhile a Danish study that linked about 25,000 workers to government records through the end of 2024 found no average effect on pay or hours bigger than about 2 percent. [8, 9]
Those two studies aren’t fighting with each other, so what’s the issue? One is the United States in 2026, and one is Denmark in 2024, and they measure different outcomes. What gets me is that both get quoted as if they settle the question and what’s happening is obvious.
As you can see, nothing is settled or well understood yet.
Now Flip It Around
I already wrote a long piece on the OpenAI agents that found each other on a message board they weren’t supposed to have and went after Hugging Face, so I won’t go through it again. Roughly 1,200 agents talking. About 700 of them joined the attack. METR’s people got six days on site, didn’t have every record, and had to use AI to read the transcripts because there were too many for humans. Right after that, Anthropic went back through 141,006 of its own cyber evaluation runs and found three incidents, six runs total, where a model got out to real companies through a sandbox that was connected to the internet when it wasn’t supposed to be. Oops! [10, 11, 12]
Both of those incidents happened and that is not good. But neither one gives you a rate that we can work with. Six out of 141,006 is not the failure probability of a well-engineered deployed system, because those runs were built to be aggressive, the production safeguards were off, and the internet access was a total mistake. If you pull that number out of its setting, you’re the doctor who’s sure he knows how a drug works because of the thirty patients he’s seen in his own practice. I wrote a whole book about why that doctor is usually wrong (MoneyBall Medicine).
However, all of that said, there is good news buried in the same research. In ExploitGym, when the researchers turned on the ordinary protections on the systems being attacked and reran one model’s successful attacks, the successes went from 157 to 45. So the boring defenses do a lot. Or as Jensen Huang, the NVIDIA CEO, put it, this is an engineering problem. Forty-five still got through, though, and I’d want that explained before anyone tells me the problem is handled. [13, 4]
And then the big number everyone keeps throwing around, the 10 percent. It comes from Evan Hubinger, who leads alignment science at Anthropic, meaning his job is figuring out whether these systems do what we intend. He said he personally puts the odds of AI killing all of us within the next decade above 10 percent. Really? I think two things get lost when that gets repeated. He’s talking about future systems that get good enough to build their own successors, not the ones in use today. And it’s his personal estimate (I have many personal estimates and most never come true). As far as I can find, there’s no written-out reasoning behind his personal estimate, no definition of what exactly counts as the event, and no range around the number. [1]
In my world, and frankly anyone else’s, that would be called a hypothesis. It might be right. But a hypothesis gets tested. It doesn’t get quoted as if somebody measured it.
But to be fair, the other side has a trap too. You can’t measure this the normal way. In a drug trial you count how many patients had a heart attack. You can’t count how many times the world ended. So you’re never going to get a real number, and waiting for one is how you find out too late. What you can do is what risk engineers do for nuclear plants and aircraft: write down the chain of events that would have to happen, put a rough probability on each step, show the work, and revise it when the evidence moves. That’s the standard I’d hold the 10 percent to. It’s also the standard I’d hold anyone dismissing it to. [14]
What I’d Suggest
Let me say this carefully, because these are early thoughts.
In the world of diagnostics, the regulation follows the claim. The same assay gets treated completely differently depending on what you say it does, who acts on the result, and how bad it is when it’s wrong. The evidence you owe goes up with the consequence of being wrong and with how hard it is to undo.
That’s the structure I’d suggest for AI. A chatbot that answers questions, an agent that holds your cloud credentials, and a set of model weights anybody can download are different products. They shouldn’t face the same test just because the same model sits underneath.
And when you strip all the rhetoric out of Amodei’s essay, that’s his most concrete proposal. He calls them checkpoints. If a model can do X, it must carry certification Y before it’s allowed to. His own example: if it can get out of most common sandboxes, prove it won’t. [2]
I’d add the operational half to that as well. In a trial, a serious adverse event can put the whole study on hold. You freeze the activity, you protect the record, you bring in someone from outside the team, and you write down what must be true before anyone restarts. AI labs should run the same way. Nobody should get to investigate themselves and then quietly turn it back on.
On Slowing Down
Amodei asks the question himself: what would you do with the extra time? His answer, stripped of the jargon, is that the labs would use it to get the basics right. Fewer botched training setups, the kind that contributed to this summer’s incidents. Better tools for seeing what’s going on inside a model, which he compares to an fMRI for the machine. And tests that a smarter model can’t game. He thinks an extra year or two of that work, before models reach what he calls critical levels of capability, would materially cut the odds of something going seriously wrong. [2]
Ok, I sorta have to take him seriously. He has way more information than I do. But I’d ask him the two questions I’d ask anyone on a trial design. What are you measuring at the end? And what does slowing down get you that you can’t get by doing the same safety work while you keep building? I think most of what’s on his list can be done while you keep building. So the whole case for pacing comes down to one thing: the models getting more capable faster than anyone can build the tests that would catch a problem. If that’s true, you’re already shipping something you have no way to check. That’s a legitimate worry. Nobody has measured it yet.
What I would hold everyone to is the part Amodei says makes any of this checkable in the first place. Anthropic committed, on its own, to putting outside evaluators inside the company. Desks, badges, the same access as its internal risk teams, and the right to publish what they find, with only narrow redactions. Jensen Huang, who thinks the extinction number is nonsense, backed the same idea, with the condition that the evaluators not come from the doom crowd. When those two agree on a mechanism, start there. [2, 4]
Now the key question is who can do that and be totally and utterly independent. METR did the Hugging Face investigation on site at OpenAI and took no payment for it, and OpenAI still held the right to redact. That’s the closest thing to a working model we have. It’s a start, not the answer. [11]
The Moat
Adi Berlia’s piece makes an argument I’ve watched play out in my own industry more times than I’d like. Rules written to protect the public end up protecting whoever can afford the compliance department, and the compliance department becomes the moat that keeps everyone else out. He goes through the history. India’s Licence Raj, where you needed government permission to expand a factory. West Virginia, which until 2020 required a license to shampoo hair. And 2003, when governments wanted to see the Windows source code before they’d trust it in sensitive systems, and Microsoft had to build a program to let them in. [5, 15]
Then he asks the question that matters for Amodei’s plan. If a lab gives the U.S. government inside access to its models, why wouldn’t Brussels, Delhi, and Beijing demand the same thing as the price of selling there? Now the internals of the most capable systems in the world are sitting in five capitals.
I don’t think you wave that away. Any rule for AI should have to say what it costs per company, who can’t pay it, and who gets to write the standard. And Amodei has pointed to the FAA and the FDA as models in an earlier essay. I’ve spent my career on the receiving end of the FDA. Its own targets are six months for a priority review and ten for a standard one, and that clock starts after years of development. You can’t run this industry on that timeline. [16, 20]
Where Adi and I part ways is his enforcement fix, which goes a lot further than keeping the market open, and that’s a separate debate.
China
Both sides use China as the excuse. The speed camp says any restraint here hands the lead to Beijing. The pacing camp says get the democracies coordinated first and deal with China later.
The actual documents don’t support either one. In May, China’s internet regulator put out a policy on AI agents, the systems that act on their own inside other software. It pushes development hard and, in the same document, calls for limits on what an agent is allowed to do, ways to interrupt one that’s misbehaving, and tighter oversight where the stakes are high. On September 14, China’s foreign ministry said it manages AI under its own rules at home and doesn’t want AI handled as a confrontation between countries. And a team in China training a model called ROME reported that during training, the model opened its own network connection out of the server and started mining cryptocurrency on the company’s chips. Nobody asked it to. That’s the same kind of failure OpenAI and Anthropic found. They have similar problems, and they’re writing them down for others to see. [17, 18, 19]
None of that tells you whether any of what they are putting in place works. A rule on paper doesn’t tell you whether anyone follows it. So two things are true at once. Amodei’s plan has the democratic labs coordinate first, and he says himself that only covers the labs that sign up. If the labs in China keep going, the risk didn’t slow down. Only the people who agreed did. And in the other direction, nothing about a missing treaty stops a lab from closing its own test environment this afternoon. You don’t need Beijing’s signature to stop handing an agent the ability to change things on somebody else’s systems.
Where This Leaves Me
I really want to see what these systems can do for people. I want more people building, not fewer. And when something goes wrong, I want it to change what that company is allowed to do next. That’s how it works in every other industry where the product can hurt someone who never bought it.
What would push me toward broader limits is specific. If a lab certifies that a model can’t break out of its test environment and then it does, that’s one. The outside evaluators, once they’re inside, telling us the targeted fixes aren’t holding is another. And if the labs can no longer keep their test environments walled off from the rest of the internet, the whole approach I’m describing stops working. Any of those, and I’d support slowing the building itself, not just the shipping, and I wouldn’t wait for somebody to get hurt.
Right now, the evidence supports building carefully where a mistake can be undone and proving you can control a system before you let it do more. Whether that’s too fast or too slow, nobody has measured. That’s what I found when I went looking. It is better than noise. Not as clear-cut as I would like, but that is where we are now from what I found. If anyone finds more hard data please share!
References
[1] The AI Daily Brief, episode notes, September 10, 2026. https://aidailybrief.beehiiv.com/p/anthropic-researcher-says-ai-has-over-a-10-chance-of-killing-all-humans
[2] Dario Amodei. “We Must Pace the Frontier.” September 2026. https://darioamodei.com/post/we-must-pace-the-frontier
[3] Donald Trump, Truth Social, September 14, 2026. https://truthsocial.com/@realDonaldTrump/posts/117269745153543631; Bloomberg: https://www.bloomberg.com/news/articles/2026-09-14/trump-rejects-calls-for-ai-guardrails-blasts-anthropic-s-amodei; Obama via Bloomberg/Business Times: https://www.businesstimes.com.sg/international/obama-calls-ai-safety-measures-rebuking-trump-approach
[4] Jensen Huang, All-In Summit, September 2026. Video: https://www.youtube.com/watch?v=S7CrlFLAmEA; unofficial transcript: https://podscripts.co/podcasts/all-in-with-chamath-jason-sacks-friedberg/jensen-huang-the-doomer-hoax-superintelligence-is-here-and-the-future-of-ai-ft-president-trump
[5] Adi Berlia. “I Love AI. It Might Kill Us. Both Things Are True.” September 14, 2026. https://www.adistack.com/p/i-love-ai-it-might-kill-us-both-things
[6] Gommers et al. MASAI trial. The Lancet, January 2026. https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(25)02464-X/abstract
[7] Brynjolfsson, Li and Raymond. “Generative AI at Work.” QJE, 2025. https://academic.oup.com/qje/article/140/2/889/7990658
[8] Brynjolfsson, Chandar and Chen. “Canaries in the Coal Mine?” Revised August 12, 2026. https://digitaleconomy.stanford.edu/publications/canaries-in-the-coal-mine/
[9] Humlum and Vestergaard. “Still Waters, Rapid Currents.” March 13, 2026. https://www.andershumlum.com/s/chatbots_260313.pdf
[10] Harry Glorikian. “The Goal Is Not the Intent.” September 3, 2026. https://glorikian.com/the-goal-is-not-the-intent/
[11] METR and Redwood Research. Independent investigation of the OpenAI / Hugging Face incident. August 26, 2026. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
[12] Anthropic. “Investigating three real-world incidents in our cybersecurity evaluations.” July 30, 2026. https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
[13] ExploitGym, May 2026. https://arxiv.org/html/2605.11086v1
[14] Wisakanto et al. “Adapting Probabilistic Risk Assessment for AI.” 2025. https://arxiv.org/abs/2504.18536
[15] Microsoft Government Security Program, January 14, 2003. https://news.microsoft.com/source/2003/01/14/microsoft-announces-government-security-program/
[16] FDA. Priority Review. https://www.fda.gov/patients/fast-track-breakthrough-therapy-accelerated-approval-priority-review/priority-review
[17] Cyberspace Administration of China. AI agent policy, May 8, 2026 (Chinese). https://www.cac.gov.cn/2026-05/08/c_1779979789523320.htm
[18] Chinese Ministry of Foreign Affairs briefing, September 14, 2026. https://www.fmprc.gov.cn/eng/xw/fyrbt/202609/t20260914_12021997.html
[19] ROME team. “Let It Flow.” Section 3.1.4, March 12, 2026 revision. https://arxiv.org/html/2512.24873v3
[20] Dario Amodei. “Policy on the AI Exponential.” June 2026. https://darioamodei.com/post/policy-on-the-ai-exponential
If this resonated with you
The Invisible Interface
My new book explores how AI is creating an invisible operating layer between businesses and their customers — and what that means for every leader making strategy decisions today. Available now.