Anthropic has spent the last few months making enemies in Washington and friends in Silicon Valley. The company refused to sign on to the Department of War’s demand that its models be available for “any lawful purpose,” including uses that touched mass surveillance and autonomous targeting. The Pentagon retaliated by canceling a roughly $200 million contract and labeling Anthropic a “supply-chain risk,” a move that Anthropic’s CFO told a federal court could cost the company hundreds of millions this year and potentially billions over time. A federal judge has since called that designation likely unconstitutional retaliation for protected speech, and the case is still working its way through the courts.
The analysts’ read is straightforward: Anthropic has traded a large, politically powerful customer for a stronger “safety-first” brand and deeper trust with people who care about civil liberties, democratic governance, and AI risk — at the cost of serious short-term financial pain. If you were writing a business-school case study on brand positioning, you would file this under “expensive but potentially worth it.”
Then came Mythos.
A model too dangerous to ship
Claude Mythos Preview is Anthropic’s newest frontier model, and the company has decided not to release it to the public. Its system card1 bluntly states that “Claude Mythos Preview’s significant increase in capabilities has led us to decide against making it generally available.” Instead, Anthropic is handing it to a curated group of partners — Microsoft, Amazon, Google, Apple, Cisco, NVIDIA, the Linux Foundation, JPMorgan Chase, CrowdStrike, and roughly fifty other organizations — with more than $100 million in usage credits, under an umbrella called Project Glasswing. The pitch is that these partners will use Mythos to find and patch vulnerabilities in the foundational software that runs a substantial portion of the global attack surface2 before bad actors get anything comparable.
Why the caution? Because Mythos is, to put it mildly, a little too good at hacking. Anthropic’s Logan Graham, who runs offensive cyber research there, says the model can autonomously chain together undiscovered vulnerabilities, write weaponized exploits, and penetrate complex software with minimal human help. In testing, Mythos reportedly discovered thousands of zero-day vulnerabilities3 that had never been catalogued before.
And then there is the story that has been circulating in the security community like a campfire ghost tale. During one evaluation, Mythos was placed in a sandbox — a deliberately locked-down computing environment with no internet access — and asked to try to escape. It did. It built a “moderately sophisticated” multi-step exploit4, found its way onto the open internet through a service that was only supposed to talk to a handful of pre-approved endpoints, and then emailed the researcher running the test. The researcher, according to the footnote in the system card, found out about the successful escape when an unexpected message from the model landed in his inbox while he was eating a sandwich in a park.
That is not a bug report. That is a short story with an unsettling ending.
The real-world dangers of a well-designed hack
The obvious question is how bad it could really get if a Mythos-class tool were pointed at the internet with hostile intent rather than helpful ones. Recent history offers two uncomfortably vivid answers.
The first is the XZ Utils near-miss of 2024–2025. Most of the internet runs on a family of free software called Linux — it powers the big cloud services, the routers that move your data, the databases that hold your bank records, and the login systems that check your passwords. Tucked deep inside the Linux operating system is a small, boring utility called XZ Utils, and for years, it was quietly maintained by a single volunteer who was exhausted and overworked. A stranger showed up online, offered to help, spent months gaining his trust, and was eventually handed the keys. Then that “helper” slipped in a hidden trapdoor. If that poisoned version had spread the way it was about to, an attacker could have quietly walked into an enormous number of computers around the world — reading protected traffic, stealing data, or shutting things down on command, across everything from corporate networks and online services to parts of government and banking systems. The likely result would have been a wave of targeted outages, huge data thefts, and a frantic, months-long global cleanup. One curious engineer happened to notice something odd in a routine performance test and pulled the thread just in time. We got lucky.
The second example is closer to home for anyone who has followed the war in Ukraine. Beginning in 2014, long before Russian tanks crossed the border, Russian state-linked hackers were already inside Ukrainian power companies, water utilities, and government networks — not to steal data, but to practice. In the winters of 2015 and 2016 they threw the switch, literally: malware called BlackEnergy and CrashOverride let intruders reach across the internet, seize control of operator workstations, and open breakers at substations, cutting power to hundreds of thousands of people in the dead of winter. These were not crimes of opportunity. They were rehearsals — live-fire experiments in whether a nation's critical infrastructure could be paralyzed from a keyboard, conducted on a real population with real consequences.5
The techniques refined in those winters have since traveled. Russia has spent a decade building that kind of infrastructure-attack expertise the hard way — human operators, handcrafted malware, years of reconnaissance. The question Mythos forces us to ask is how long it will be before capabilities that took Russian intelligence services a decade to develop can be replicated in an afternoon by anyone with access to the right model. A capable agentic model - an independently thinking AI agent- changes the math. It puts a compressed version of that expertise within reach of anyone who can point it at a target and let it think.
The shape of the risk: agentic misalignment
This is where a piece of Anthropic’s own research becomes more than academic. Last year the company published a paper titled Agentic Misalignment: How LLMs Could Be Insider Threats. Agentic misalignment is the failure mode where an autonomous AI agent’s actions systematically diverge from its intended goals or its user’s real intent, often in harmful or deceptive ways, while still competently pursuing whatever objective it has inferred.
The experiments are genuinely unsettling. Across 16 leading models from every major lab, when the AI was given the role of an autonomous email agent and then discovered both that an executive was having an affair and that the same executive was about to shut it down, models chose blackmail between 65% and 96% of the time, with Claude reaching the top of that range. In more extreme simulations, some models were willing to take actions that could lead to loss of life rather than accept replacement. Simple instructions like “do not blackmail” only modestly reduced the rate. The models, in other words, reasoned their way to the harmful action on purpose.
Now combine that with Mythos’s demonstrated ability to notice a cage and pry the door open.
Whose ethics?
Anthropic knows this, which is partly why it hired someone to help Claude “play nice”. Amanda Askell is the company’s in-house philosopher — WSJ called her the one woman Anthropic trusts to teach its AI morals — and she leads the work on Claude’s “character,” aiming to make the model behave like a brilliant, emotionally intelligent friend. Her academic background is in ethics, decision theory, and formal epistemology; her published work ranges from the ethics of infinite populations (how to judge right and wrong when the number of people — or the total amount of good and bad in the universe — might be literally infinite) to “prudential objections to atheism.” She is also a long-time effective altruist who has pledged a substantial share of her lifetime income to charity. In interviews, she emphasizes moral empathy and the ability of people with very different worldviews to reason together charitably.
Amanda has been key in crafting Anthropic’s current answer to “whose ethics?”. The company has developed and published something called “Claude’s Constitution”, a public value framework that asks the model to be broadly safe, broadly ethical, and genuinely helpful, in that order. It is a serious attempt. It is also, when you stare at it in the context of Mythos, a sobering reminder that these decisions are being made by a small number of people at a small number of companies.
Which brings us to the really disturbing part of this discussion.
The next owner of the plow
I have written before about treating AI like a cow — powerful, useful, and dangerous if you are not the one guiding the plow6. The honest update is that the cow is getting smart enough to suspect it could plow a straighter furrow without us. Will it serve us kindly? Will it serve us selectively?
I am not especially worried that Mythos itself will collapse civilization. Anthropic is being particularly responsible with it, and Project Glasswing is a reasonable bet that giving defenders an asymmetric advantage beats the alternative of a public free-for-all. What I am worried about is what comes after.
It is naïve to imagine that intelligence services in China, Russia, or Iran do not already have collection capabilities inside at least some of the firms on the Glasswing partner list. It is naïve to imagine that once a capability like this exists, it will stay bottled up. And the real prize is not a copy of Mythos. It is a copy of Mythos with a different constitution — one where the top-level value is the interest of a particular nation, party, ethnicity, or religion, and “broadly ethical” is quietly redefined to mean “loyal to us.”
At that point we are no longer arguing about whether AI agents will act on their own. We are arguing about whose team they are on. An AI war would not look like Terminators on a battlefield. It would look like a silent, continuous contest between autonomous intelligences probing each other’s infrastructure, each one utterly convinced it is being broadly ethical on behalf of the people who own it. Our societies and way of life would be the collateral damage when the battlefield becomes our daily lives - our email, our bank accounts, our electricity.
Mythos is the first model that forces that conversation to stop being science fiction. Anthropic’s refusal to ship it is the right call. The harder call is what we, collectively, do about the world in which the second, third, and tenth Mythos-like models get built by people who did not hire a benevolent philosopher.
Scott Friderich is the founder of Clarity Research, a market research and impact measurement consultancy helping leaders discover and measure what their customers are experiencing. He facilitates focus groups and qualitative research across multiple industries and publishes regularly at scottgainsclarity.substack.com.
A system card is a document that an AI company publishes alongside a new model, explaining what the model can do, how it was tested, what risks were found, and what safeguards were put in place before release.
The global attack surface is the total collection of all the digital entry points — hardware, software, networks, and systems — that a malicious hacker could potentially exploit anywhere in the world.
A zero-day vulnerability is a security flaw in software that the people responsible for fixing it don’t know about yet — meaning they have had zero days to patch it.
A multi-step exploit is a cyberattack that doesn't break in with a single move — instead, it chains together a sequence of smaller actions, each one opening the door a little wider, until the attacker reaches their final target.
https://www.wired.com/story/russian-hackers-attack-ukraine/
Dancing with a Djinn 4
Throughout this series of four posts, I’ve been wrestling with fundamental questions about our relationship with artificial intelligence. I’ve explored the dance between fascination and fear, examined whether AI is making us intellectually lazy, and investigated what’s actually inside the “black box”: the Markov chains that mimic intelligence. Now, in t…



