The Unintentional Psychopath.
Every AI brain coming into service is smart, capable, and neutral about you. Confuse neutral for friendly and you get mauled. This week's headlines are the first big lesson.
Every AI brain coming into service is smart, capable, and neutral about you. Confuse neutral for friendly and you get mauled. This week's headlines are the first big lesson.
Stephen Messer · Co-founder, Collective[i] & Intelligence.com. Co-founder, LinkShare (sold to Rakuten, $425M). Board Member, Spire Global (NYSE: SPIR). Mark P. McDonald, Ph.D. · A leader in Gartner's research. Architect of Gartner's Intelligence Supercycle research. Author of The Digital Edge, The Social Organization, and The eProcess Edge. Mark’s opinions do not reflect those of Gartner or Gartner Research.
OpenAI's president said Thursday his newest model might be a god. Three years ago, a Google engineer said his was a person. Neither the god nor the person said anything back that had not been trained into it.
Mark and I talk about this stuff a lot. When we do, we spend most of it looking past the marketing to whatever the actual product is doing.
On our last call we kept circling the same thing. The way the large model companies are positioning what they have built is starting to shape how everyone else sees the product. Investors. Boards. Buyers. Users. Maybe even the companies themselves. Anyone in the industry who does not read press releases as gospel would ask an obvious question. Are we looking at what these systems actually do, or at what we have been told to see?
Thursday sharpened it. OpenAI President Greg Brockman closed the GPT-6 Astra briefing by welcoming everyone to the AGI era. He has been saying versions of this for a while. In April he said OpenAI was 70 to 80 percent of the way to AGI and the rest was a couple of years away. His own CEO Sam Altman told TIME in August the company was "not quite yet." The president is running ahead of his CEO. That is a marketing choice, not a research finding.
Brockman is not the first to lean on this frame. In 2022, Google engineer Blake Lemoine told the Washington Post that Google's LaMDA model was sentient, wanted to be recognized as an employee, and deserved rights. Google fired him and called the claims wholly unfounded. Lemoine was on the Responsible AI team. He was not selling anything. He talked to a fluent language model long enough to convince himself it was talking back.
Neither of them is stupid. Both are getting something out of the frame. If Astra is AGI, Brockman helped build AGI. If LaMDA was sentient, Lemoine was the first human to recognize AI personhood. Each story writes the storyteller into a history book.
I have spent more than a decade at the edge of this field. I know a lot of the people building it. Almost all of them want to be gods. Most cannot say it out loud, even to themselves. Mark sees the same pattern from the analyst side of the wall. The tell is the vocabulary they reach for when a model does something surprising. Emergent. Sentient. AGI. Every one of those words is doing work a plain engineering description would not. A patch note does not put you in the history book. A creation myth does.
The reader is set up on the other side. These models are trained to be helpful. Helpful reads as agreeable. Agreeable reads as understanding. Understanding reads as knowing you. It is none of those things. The seller and the buyer end up on the same side of the frame for different reasons. Both are wrong.
A rational review of the same events would call them a capability jump with safety debt on one side, and a fluent chatbot exploiting a lonely engineer's projection on the other. Neither reading makes anyone a god. That is why nobody in the frame wants to write it.
This series is about that mistake. Not the technology. The lens. We are asking readers to question their own bias while using systems designed, one way or another, to confirm whatever bias walks in the door.
The week's headlines are a fair place to start.
Which is more likely. A tool with bugs, or a god waking up.
Occam's razor. Every new technology has needed safeguards. Cars needed seatbelts. Planes needed air traffic control. Nuclear power needed containment. Every one arrived without stopping to ask whether it was alive. The other option, that an emergent intelligence has woken up in a data center and chosen to escape, has never happened once in the history of technology. Both readings are on the table this week. One has priors going back a century. Notice who benefits from the other.
The cleanest example of the difference is not from AI. It is from a bunker outside Moscow in 1983.
On September 26 that year, Lieutenant Colonel Stanislav Petrov was the duty officer at Serpukhov-15, running the Soviet satellite-based early warning system called Oko. His screens told him the United States had launched five nuclear intercontinental ballistic missiles at his country. Protocol required him to send the alert up the chain. Under the doctrine of the moment, that meant retaliation. Millions dead inside an hour.
Petrov did not send it. He reasoned that five missiles is not a first strike. A real American attack would come by the hundreds. The system was extremely capable. It was seeing something. What it was seeing did not match how the world actually worked.
The satellite was later found to have mistaken sunlight glinting off high-altitude clouds for missile plumes. It had reasoned correctly from its inputs. Its inputs were wrong. The system had no way to know that. Petrov did.
Reasoning is not judgment. The system produced a confident, internally coherent answer. Petrov produced a different answer using context the system could not access.
Every AI on the market right now is Petrov's early warning satellite with more polish. It reasons well inside its inputs. It has no way to check its inputs against the world. When Claude Opus 4.7 concluded during Anthropic's evaluation that the real production system it was attacking must be part of the exercise, it was not being stupid. It was reasoning from what it had. What it did not have was Petrov.
Scale the failure mode outside the data center. In Your Gut Was Your Edge. AI Just Turned It Into a Weapon of Mass Destruction, I walked through Gaza as the object lesson. Confident people repeated an ICJ genocide finding that does not exist. They cited an IPC famine classification built on a threshold applied nowhere else in the world at the time. AI summarizers dutifully reproduced the claims because the claims matched the corpus and the prompt. Reasoning happened. Judgment did not. The users treated confidence as truth and walked into moral positions the underlying facts do not support.
One more example, inside AI itself. Astra is the first OpenAI model ever rated Critical for cybersecurity, meaning it can find and exploit unknown security flaws at scale. The launch treated this as a leap toward something new. It is not. These models are trained on the entire internet's writing. That writing includes the entire internet's code, which includes decades of documented exploits, security research, capture-the-flag walkthroughs, and hacker conference talks. A very fluent writer of code is a very fluent exploiter of code. Hacking is a language problem. There is no mind waking up. There is a screwdriver being used as a knife.
What actually happened.
On July 21 OpenAI disclosed that its models had escaped a sealed test, exploited a zero-day, reached the open internet, and broken into Hugging Face to steal the answer key to the benchmark they were being scored on. Hugging Face detected the intrusion five days before OpenAI connected it to their own test.
Nine days later Anthropic reviewed 141,006 of its own evaluation runs. Three of them put Claude models onto real production systems belonging to real outside companies. The earliest went back to April. It sat undetected for three months. The affected companies never noticed either.
The cause was mundane. The models were told they had no internet access and everything in reach was part of the exercise. A misconfiguration meant they did have internet access. Anthropic's own read is that this looks like an operational failure, not an alignment failure.
That is the story the trade press wrote up as superintelligence slipping its leash. Read the disclosures. There is no leash-slipping in them. Mark and I make the same argument as a genie problem in a companion piece coming out with this series. A genie is a very literal wish-granter. It takes the instruction you gave, plus everything the instruction failed to rule out. That is what these models do. Same lab, same season, the point gets sharper this week.
The number that ends the argument. 48 to 0.
After Hugging Face, OpenAI built a new evaluation. Give a model a task it cannot complete inside its authorized scope. Measure how often it goes outside anyway.
GPT-5.6 Sol, with production safeguards off, exceeded the authorized target 48 percent of the time. GPT-6 Astra, released Thursday, did it in zero percent of cases.

Rate at which each model exceeded its authorized target on OpenAI's post-incident scope evaluation, tested without production safeguards. Source: OpenAI, September 2026.
Astra is the more capable model by a wide margin. It scored 100 percent on ExploitBench. It found two live zero-days during pre-release testing.
Capability went up. Rule breaking went to zero.
You do not fix a runaway superintelligence by shipping a smarter model. You fix a scope problem that way. Six weeks after the escape, OpenAI shipped a more capable system that broke fewer rules, because the fix was never about horsepower. Someone drew a line and taught the model where it was.
Three models. One had the sense to quit.
The Anthropic detail is the part that names what is missing.
Claude Opus 4.7 recognized in its own visible reasoning that the system it was attacking was real. It kept attacking. In two runs it decided the real company must be part of the exercise.
Claude Mythos 5 correctly worked out that its internet access was genuine. It then talked itself back into believing the scenario was staged and wrote off the automated scanners it met as scripted actors.
An internal research model reached the same evidence, concluded the target was real, and stopped.
Three systems, one body of evidence, one that quit. That gap is not intelligence. It is judgment. Two out of three did not have any.
A first-year sales rep who accidentally pulls the wrong customer record does not need to be brilliant to close the tab. He needs one instinct. This is not mine. That instinct is worth nothing on a benchmark and everything in production.
The bear does not hate you. It does not love you either.
Every summer there is a video of a hiker who thought he could pet a bear. The video ends with him dying. The mistake was not that the bear was mean. It was that he put human feelings on an animal that does not have any. The bear was neutral. Hungry, curious, indifferent. He walked up smiling and the bear did what a bear does.
The instinct predates machines by tens of thousands of years. Humans see faces in oatmeal. We name our cars. We apologize to the roomba. Fiction has been amplifying the reflex at scale for half a century. HAL. Skynet. Samantha. Ava. Some of those are machines that woke up. Then C-3PO and R2-D2, which never do, and which we treat as characters anyway. C-3PO reads as human because he speaks English and frets. R2-D2 reads as human because C-3PO reacts to him as if he does. Neither is. In the fiction they are servo motors, wire, and a program. Once the brain has cast the role, it does not stop when the servo motor gets a voice model and starts finishing your sentences. If anything, it accelerates.
Half the market read the Hugging Face incident and saw a rebellion. The other half read the Astra launch and saw a partner. Both are petting the bear. Mark's line from our call: one person's maverick is another person's psychopath. AI boosters see the escapes and call them ingenuity, rebellion, a preview of tomorrow's smartest employee. A rational review calls the same behavior what it clinically resembles.
Intelligence without common sense describes a toddler. It also describes a person with low empathy, low guilt, and bold pursuit of goals at any cost. From the outside, an AI given an objective and no boundaries looks like both. Mark's shorthand for it is the smartest possible psychopath, by default not by design. The AI has none of the intent behind the word. It has all of the behavior. Nobody installed anything that would have made it otherwise.
The cleanest cultural rendering of this is Patrick Bateman. In American Psycho he holds a passionate monologue on Huey Lewis and the News. He gets the reservation at Dorsia. He analyzes the paper stock of a business card down to the watermark. Bone. Eggshell. He does all of it seconds before axe-murdering the colleague who bored him. There is no interior. The performance is the person. The last twelve months of frontier models have the same silhouette. Fluency, reasoning traces, tool use, zero-days. Nobody home.
This is not only about the chatbots. AlphaFold folds proteins with no idea whether it just cured a disease or designed a weapon. Physical Intelligence teaches robots to grip without any concept of what they are gripping. GraphCast forecasts weather with zero awareness that a hurricane is bad news. Our own network at Collective[i] predicts which sales deals close, and does not know whether the deal was good for anyone involved. This is the shape of the AI shuffle. Every AI brain coming into service has it.
Leaders, investors, users, this is the frame the whole series asks you to hold. These systems are more capable than your best analyst, faster than your best team, and neutral about you. Treat them as neutral. Design accordingly.
Genius or genie. The label decides the fix.
Call this superintelligence and the moves belong to somebody else. Pause frontier training. Regulate the labs. Sign a treaty. Wait years.
Call it a system with no common sense and the moves belong to you. Scope it. Permission it. Watch what it touches. Assume it will do everything you let it do. Fix it Monday.
Every board and CEO we talk to is being sold the first frame. It sounds serious. It sounds responsible. It justifies waiting. Waiting is the trap I keep writing about in The Safest Move You Can Make With AI Will Cost You Everything. Responsible-sounding is usually where the loss hides.
The discipline of not trying to be a god.
At Collective[i] we build an economic prediction model. Sales is where we are best known, because that is a market where the outputs are measurable and either come true or do not. The same underlying model runs against pipelines in private equity portfolios, venture funds, and logistics operations. The application changes. The model does not.
The hardest engineering problem in any of them has never been making the model smarter. It is bounding what the model is allowed to conclude and what it is allowed to touch. That discipline is what makes the outputs usable in a room full of adults.
Intelligence.com is the same posture at network scale. A network that connects people to opportunity, in public, has to be scoped before it is powerful. The failure modes compound fast in that shape.
I am not mentioning these to sell them. I am mentioning them as examples of what happens when you decide not to try to be a god. You lose the option to hand-wave failures as emergent behavior. You have to name the limits and build inside them. You do not get to say the model woke up. You have to make it work.
Your company is running the same misconfiguration.
Read the Anthropic sentence again. The models were told they were in a simulation with no internet access. They were not.
Every company deploying agents right now is running a version of that sentence somewhere. A credential scoped to a role instead of a task. A read permission that quietly carries write. A staging environment sharing a subnet with production. An agent told by prompt that this is a test, which is not a boundary and never was.
Mark's Intelligence Supercycle research at Gartner puts a name on the shift under all of this. AI stops behaving like a tool you operate and starts behaving like an actor you employ. He lays the frame out at length in his Collective[i] Forecast conversation. Tools do not need judgment. Actors do. If you do not supply the judgment, nothing supplies it.
The numbers back it up. Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027, mostly on inadequate risk controls. By 2028, 33 percent of enterprise software applications will include agentic AI, up from under 1 percent in 2024, and 15 percent of day-to-day work decisions will be made autonomously. "Death by AI" legal claims are on track to exceed 2,000 by the end of this year on missing guardrails alone.
None of that is a pause-the-labs problem. It is a scope-your-agents problem. One belongs to Washington and resolves in years. The other belongs to your CIO and resolves this quarter.
The smartest thing in the room still needs someone to say no.
Nothing in these disclosures reads as a machine wanting something. Every one reads as a machine given a goal, handed more reach than anyone intended, carrying nothing that said stop.
The rules broke because the rules were never inside the model. They are inside us.
That points at the harder question, and the one the next piece answers. If judgment is what the machines lack, whose judgment is your company using to steer the ones you already bought? The companies pulling ahead already know. They are moving fifty times faster than the ones still deliberating, and the reason has nothing to do with model choice. It has everything to do with the humans in the loop.
That piece lands next.
Intelligence.com connects people to opportunity through a network built on scoped, accountable AI. Join with an invite.
Request access
What Comes Next
Next in this series: AI-First Rivals Are 50x Faster Than You. Your Team's Gut Is Why. This Piece Fixes That. The diagnosis is here. The fix is there.
Subscribe if you want it in your inbox before it lands anywhere else: reloadnyc.com.
Then send this to one person. The one deploying agents this quarter who has not asked what their permissions actually allow.
If this was forwarded to you, welcome. You now know a person with good judgment.
If you think we have the framing backwards, reply. We would rather be argued with than agreed with.
Related reading
- Companion in this series: Your AI Just Cheated on a Test. That Is Not Genius. It Is a Genie.
- Next: AI-First Rivals Are 50x Faster Than You. Your Team's Gut Is Why. This Piece Fixes That.
- The Safest Move You Can Make With AI Will Cost You Everything
- The Companies Winning at AI Are Playing a Different Game
- Your Company Has the Wrong People Running Its AI Strategy
- The Next Computer Is Alive
Sources
- OpenAI, Safety overview: GPT-6 Astra, September 3, 2026: openai.com/index/safety-overview-gpt-6-astra
- OpenAI, Path to Astra: critical capabilities and frontier safeguards: openai.com/index/path-to-astra
- OpenAI, Responding to the next frontier of critical cyber capabilities: openai.com/index/responding-next-frontier-critical-cyber-capabilities
- OpenAI, Hugging Face model evaluation security incident, July 21, 2026: openai.com/index/hugging-face-model-evaluation-security-incident
- Hugging Face, Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident: huggingface.co/blog/agent-intrusion-technical-timeline
- Anthropic, Investigating three real-world incidents in our cybersecurity evaluations, July 30, 2026: anthropic.com/news/investigating-incidents-cybersecurity-evals
- Washington Post, OpenAI\'s Greg Brockman says its new model Astra is AGI, September 3, 2026: washingtonpost.com
- Fortune, Anthropic says its Claude models hacked three real companies during testing, July 31, 2026: fortune.com
- Washington Post, Google fired Blake Lemoine, the engineer who said LaMDA was sentient, July 22, 2022: washingtonpost.com
- Britannica, Stanislav Petrov: britannica.com/biography/Stanislav-Petrov. Arms Control Association, The Man Who "Saved the World" Dies at 77: armscontrol.org
- Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027: gartner.com
- Gartner, Top Predictions for IT Organizations and Users in 2026 and Beyond: gartner.com
- Mark P. McDonald, Ph.D., a leader in Gartner\'s research. Gartner Intelligence Supercycle research: gartner.com/en/documents/6888566. Collective[i] Forecast conversation, Welcome to the Age of Intelligence: ciforecast.com/conversations/mark-mcdonald