
Most people who leave a high-profile AI lab sign the paperwork, collect the payout and stay quiet. In 2024, Daniel Kokotajlo did the opposite. Offered roughly $2 million in vested equity at OpenAI in exchange for signing a standard non-disparagement agreement, he walked away from the money so he could say publicly what he actually believed. In a July 2026 interview on The Diary of a CEO, he put it in blunt terms: a roughly 70% chance that advanced AI “goes horribly wrong.”
It is worth being precise about that number, because the precision is part of why he is credible. Kokotajlo does not claim a 70% chance of human extinction specifically. He describes a 70% chance of a catastrophe on the scale of human extinction — a broad category that also includes AI systems seizing control or slipping beyond meaningful human oversight, of which outright extinction is only one possibility.
A claim like that is easy to dismiss as another dramatic prediction in a field that produces them constantly. What makes Kokotajlo harder to wave off is his track record. In 2021, a full year before ChatGPT existed, he published a scenario document titled “What 2026 Looks Like”, describing AI chatbots woven into daily life, AI agents taking instructions inside workplace chat tools, and rising public unease about it all. Read today, large parts of it read less like speculation and more like a diary written in advance. For business leaders deploying AI systems right now, that combination — a costly personal decision plus a documented forecasting record — is worth five minutes of attention, even for those who find the extinction framing overstated.
A $2 Million Decision to Speak Freely
Non-disparagement agreements are a routine part of departures from major AI labs, and the equity attached to them is typically substantial enough that few employees decline. Kokotajlo did. By refusing to sign, he gave up roughly $2 million so that he could criticize the industry he had just left without contractual restriction. That decision alone does not make his 70% estimate correct, but it does establish that he had a strong financial incentive to stay quiet and chose not to take it — a detail that separates his warning from commentary made by people with less at stake.
The Prediction That Makes Him Hard to Dismiss
Kokotajlo’s 2021 document, written before large language model chatbots were part of mainstream life, anticipated several specific developments: conversational AI becoming embedded in everyday tools, AI agents operating inside workplace communication platforms, and a growing public discomfort with both. Reporting on the document, including coverage from The New York Times, has noted how closely its scenarios track what actually unfolded; later reviewers found that more than half of his concrete near-term predictions resolved as essentially correct. Forecasting accuracy in one document does not guarantee accuracy in the next, but it is the reason his current warnings are being taken seriously inside the AI research community rather than dismissed as speculation.
The Core Warning: Recursive Self-Improvement
The center of Kokotajlo’s concern is a specific mechanism, not a vague fear of AI in general. Leading labs, he argues, are no longer only hiring human researchers — they are building AI systems specifically designed to automate AI research itself. If that effort succeeds, progress stops moving along a roughly linear path and starts compounding, accelerating faster than existing human oversight structures can realistically track. He describes this as the industry’s “scary open secret”: when he raises the 2027–2028 timeline with insiders at OpenAI and Anthropic, he reports that they largely do not push back. His own median for superintelligence now sits around 2029 — before the end of the decade.
What “Plan A” Looks Like
According to Kokotajlo’s account — laid out in his follow-up report “AI 2040: Plan A”, published days before the interview — the industry’s default approach is to regulate AI development at the last responsible moment while still aiming to reach superintelligence deliberately, on a timeline stretching to roughly 2040. It is a plan that assumes regulation can be applied precisely and in time, an assumption his own forecasting history gives him standing to question. His earlier scenario, AI 2027, maps the faster version of that trajectory.
Expert Perspective: The Version of This Risk Businesses Face Today
At ZTS Infotech, we build AI systems for real, paying clients, which means we deal with a smaller, more immediate version of the exact problem Kokotajlo describes at the frontier. The extinction-level scenario he outlines concerns superintelligent systems operating beyond meaningful human oversight. The version that shows up in a client’s production environment is far more mundane, and far more avoidable: a line of AI-generated code that ships without a human reviewing it.
That is not extinction risk. It is a security flaw nobody caught before deployment, a financial calculation nobody double-checked, a module that quietly does the wrong thing in an edge case no one tested. But the structural problem is the same one Kokotajlo is asking governments to address at a civilizational scale: systems operating on the assumption that the AI got it right, without a verification step to confirm it. The practical response scales down cleanly. Where he is asking regulators to check AI development before it outpaces oversight, we ask developers and AI managers to check every pull request, every generated module, every piece of AI output before it ships — the same hidden human layer that reliable agentic deployments depend on.
This is not a policy debate reserved for labs pursuing superintelligence. It is an operating discipline for any team shipping AI-assisted code today. At ZTS Infotech, every AI-generated output is reviewed by a person before it touches a client system — no exceptions, and a core part of our custom AI development services. Businesses that treat human review as optional because the output “looks right” are running a smaller-scale version of the exact control problem the industry’s own researchers are now walking away from paychecks to warn about.
Key Takeaways
- Daniel Kokotajlo left OpenAI in 2024, forfeiting roughly $2 million in equity rather than sign a non-disparagement agreement.
- On The Diary of a CEO in July 2026, he estimated a ~70% chance that advanced AI “goes horribly wrong” — a catastrophe on the scale of human extinction that, he clarifies, also includes AI takeover and loss of control, not extinction alone.
- His 2021 document, “What 2026 Looks Like,” predicted AI chatbots in daily life and AI agents in workplace tools — a year before ChatGPT existed — and reviewers later judged much of it accurate.
- His central warning concerns recursive self-improvement: AI systems built specifically to automate AI research, which could accelerate progress beyond human oversight.
- He reports that insiders at OpenAI and Anthropic largely do not dispute a 2027–2028 timeline when he raises it; his own median for superintelligence is around 2029.
- The industry’s stated default approach, per Kokotajlo, is to regulate at the “last responsible moment” while still pursuing superintelligence deliberately by around 2040 — the subject of his report “AI 2040: Plan A.”
- The practical takeaway for businesses is smaller in scale but structurally identical: every AI-generated output should be reviewed by a human before it ships.
Looking Ahead
Reasonable people can disagree with a 70% probability estimate on a question this uncertain, and many AI researchers do. What is harder to dispute is the underlying discipline Kokotajlo is calling for: do not assume an AI system got it right simply because the output looks plausible. That principle does not require agreeing with his timeline or his odds — it only requires acknowledging that unchecked AI output, at any scale, is where problems start. As frontier labs and regulators work through the larger version of this question over the next few years, businesses that build human review into their own AI workflows now will be the ones least exposed when the industry’s answer finally arrives. Talk to our team about responsible AI adoption.
-
Writen by Anirban Das
USA:
India: