I’ve been involved in AI security since 2010 when I built a probabilistic graphical model of military cybersecurity risk at a think tank in LA. Nobody really cared. I was getting paid $85k a year and working 60 hour weeks and felt (and was!) extremely lucky. Models like mine were curiosities back then, alongside a dozen other throw-away DARPA-like bets in the academic research ecosystem. So it’s very strange to live in a world where AI and cybersecurity is discussed by heads of state. But the discussions appear to have gone off the rails these past few weeks and we’re now, for the first time, going in a direction that will cause real self-inflicted harm due to misconceptions about what new capabilities represent and how they can actually keep us safe.
Restricting frontier AI cybersecurity capabilities continues to be the wrong approach
Chinese models like GLM-5.2 exist now, which have near frontier capabilities, are open weight (and thus can be used on private inference infrastructure), and are good at cyber attacks, and yet somehow substantial parts of the AI safety and AI policy community believe it’s a good idea to delay and/or debilitatingly guardrail US model launches under the assumption that attackers will benefit from them more than defenders. Attackers have every incentive not to depend on these US models, which benefit defenders more.
This is because, if attackers use the American models, they’ll be monitored by teams that have relationships with law enforcement and the NSA. If they use the Chinese models, they can operate privately, fine tune new capabilities into the models, and fine tune away annoying safeguards. Cyber defenders, on the other hand, would greatly benefit from full access to the US models (when I use Fable to build my startup’s defensive cybersecurity application, I get guardrailed and downgraded to Opus 90% of the time). Somehow US policy has landed us here.
When I was at Meta interacting with the US, UK, and EU AI safety and security policy communities around LLM launches was part of my job; similar to counterparts at the other labs. What I observed (and participated in) was a capabilities-centric view of AI safety. Measuring dual-use cyber capabilities was the focus of any given model’s pre-launch evaluation. If a model exceeded dual-use cyber capability thresholds it was deemed dangerous.
Regardless of the exact politics of the Trump administration, this capability-centric framing is what got us to the policy failures of the past few weeks. What we need is an ecosystem-centric view, which would lead policymakers to recognize that in a world of uncontrolled AI diffusion in which there are 6-7 companies at or near the frontier competing to be first (some of which are committed to open weights approaches) AI cybersecurity safety is really about ensuring defenders win the race to adopting the most powerful dual-use AI cybersecurity capabilities.
We should be leaning into ensuring defenders can easily access all of Claude Fable and GPT-5.6’s capabilities as fast as possible, ensuring that defenders walk our adoption curve faster than attackers.
‘Distillation’ and expert trajectory theft
Policy insiders and the American labs themselves document distillation campaigns against their models, and there’s a grey market for access to US model inference in China (where access to these models is disallowed). Distillation, geopolitically, means something very different from what Geoff Hinton & co-authors meant in the 2010s when they proposed distillation as literally training a student model on the logits or probability outputs (sometimes in addition to hidden states) of a teacher model. Now it can mean anything from illegally using a teacher model as a judge LLM for a student LLM to using the partial trajectories these models expose as — in some form — supervised fine-tuning data.
Look at how expensive top tier training data are for frontier model capabilities and how inconvenient or prohibitive it would be to pay for large quantities of these data (below is a screenshot from Mercor’s landing page - Mercor is a leading data annotation firm).
Lucky for some geopolitical competitors to the American labs, these expert data can be ‘distilled’ from models once they’ve trained on them. And you don’t need that much of it relative to the massive scale of pretraining datasets.
Academics have famously documented that post-training on even a few thousand trajectories can help replicate the behavior of frontier models: LIMA fine-tuned a 65B LLaMa model on only 1,000 curated examples and still achieved remarkably strong performance, and a 2026 study (PassNet) distilled under 4,000 expert trajectories from Claude-Sonnet-4.6 into a small open model, nearly closing the gap to frontier performance.
The risk of distillation is built into the way models are productized. Closed model companies need to serve many trillions of tokens per year to generate an exponentially growing mass of revenue.
Despite efforts to detect distillation in these token streams, distillers will continue to succeed in extracting enough behavior to give their model building efforts a substantial boost. Why pay Mercor domain experts $100/hr to generate millions of tokens when you can distill these capabilities from finished models that are 6 months ahead of you?
Agent-to-agent contracts for security automation
My co-founders and I soft-launched our company this week. We’re working on doing what we believe is the most urgent thing today, burning down the mountain of security tech debt most organizations are stuck with, by reimagining vulnerability and exposure management agentically. As attackers adopt AI, we believe (and empirical data proves!) they’ll prioritize exploiting this latent security debt, as they already do manually, with agents.
We’re now ready to ship v1 and are meeting with prospective new partners at Blackhat. If you’re in a security leadership position and are interested in meeting in person there, or over Zoom, please reach out!
Where are all the AI cyber-attacks?
I’ve been doing some yet to be completed research on attacker demographics and how they bear on the rate we can expect various attacker constituencies to adopt AI agents. In the meantime, it’s useful to consider these ethnographic snippets (which I generated with Opus-4.8-Max on research mode) about attacker organizations when thinking about for whom guardrailed deployments Fable / GPT-5.6 might possibly benefit [very few of them, given their other options, I think!].
Criminal organisations
Scattered Spider (also tracked as Octo Tempest, UNC3944, Muddled Libra)
Young, mostly English-speaking, and native to Discord and gaming servers rather than any underground forum — investigators trace its roots to the same corners of Roblox and Minecraft communities that also produced SIM-swapping crews. There is no boss, no data-leak site, and by most accounts constant infighting; some members are minors, complicating prosecution. Its signature move is almost pre-digital: call an IT help desk, impersonate an employee found on LinkedIn, and talk a human into resetting a password. That took down MGM Resorts and Caesars Entertainment in 2023, and has since netted an estimated $66m or more in extortion, with one Mandiant executive saying he has personally seen client payments reach eight figures. AI barely features in its tradecraft; the vulnerability it exploits is organisational trust, not code.
Conti (ransomware-as-a-service; ~350 members; an estimated $2.7bn extorted, 2020–22)
Before it collapsed, Conti ran like a mid-sized tech company: a dedicated HR department, salaried coders, performance reviews and an employee of the month. Recruiters used legitimate headhunting sites, telling applicants they were hiring for an advertising firm; a leader known by the handle Stern set wage budgets and fined coders who missed deadlines. The enterprise unravelled in February 2022, when a member enraged by the group’s pro-Russia statement on the Ukraine invasion leaked 60,000 of its internal chats. None of that required AI. It required an org chart.
LockBit (ransomware-as-a-service; 2,500+ victims across 120 countries; $500m+ extorted)
For five years the group’s administrator operated purely as “LockBitSupp,” a bravura persona who ran tattoo contests, taunted researchers, and offered $10m to anyone who could unmask him. When the US, UK and Australia did exactly that in 2024 — naming Russian national Dmitry Khoroshev — he reportedly asked investigators to “give me the names of my enemies” in exchange for cooperation. LockBit had publicly apologised after an affiliate hit a children’s hospital and claimed to have expelled them — while, investigators later found, quietly keeping that same affiliate active for another year.
ALPHV / BlackCat ($22m Change Healthcare ransom, withheld from its own affiliate)
The cleanest illustration that there is no honour among cyber-thieves. After a 2024 attack on Change Healthcare reportedly paid out $22m, the group’s administrators changed their status to “GG” — good game — froze the affiliate’s account, and posted a fake law-enforcement seizure notice, recycled from a genuine one months earlier, to cover their tracks. The cheated affiliate, who still held the stolen data, resurfaced weeks later running a new extortion operation against the same victim. Also not an AI story — a fight over a bag of money.
Black Basta (500+ victims; $107m+ in traced ransom payments 2022–24; internal chats leaked February 2025)
Black Basta is the only criminal group whose internal AI usage is on record at scale. Leaked Matrix chats spanning September 2023 to September 2024 — the first full year of widely available AI tools — show operators using ChatGPT to draft phishing letters in polished English, rewrite malware code from C# to Python, debug exploits and look up target financials. The group’s leader reportedly bought ChatGPT accounts from underground marketplaces and shared credentials with the team. None of this is a capability transformation: all of it replaces work a skilled developer or a hired contractor could have done in 2021. The chats also reveal a group in slow decline — internal squabbling over targeting Russian banks led to the leak itself, in the same pattern that destroyed Conti two years earlier. The gang has been largely dormant since early 2025.
State-sponsored groups
APT29 — Cozy Bear (Russia / SVR; also Midnight Blizzard, NOBELIUM, The Dukes; active since 2008)
Russia’s foreign intelligence service runs its premier cyber unit with a patience that makes the ransomware groups look impulsive. APT29 sat inside SolarWinds’ build pipeline for months before anyone noticed in 2020, compromising 18,000 clients and reaching into US Treasury, Justice and the State Department. In January 2024 it breached Microsoft’s corporate environment via a legacy test account that lacked multi-factor authentication — then, once in, extracted email from senior leadership and cybersecurity staff to learn what the company knew about the intrusion. Since 2022 the group has pivoted away from on-premises exploits towards cloud-identity abuse: OAuth application manipulation, token forgery and residential-proxy networks that make traffic indistinguishable from legitimate users. AI use: minimal confirmed evidence. The tradecraft depends on stealth and patience, qualities that current AI systems generate at the cost of tell-tale artefacts.
APT44 — Sandworm (Russia / GRU; also Seashell Blizzard, Voodoo Bear; Unit 74455; upgraded to APT44 by Mandiant in 2024)
If APT29 is the scalpel, Sandworm is the wrecking ball. GRU’s Unit 74455 caused the two Ukrainian power-grid blackouts in 2015 and 2016, then released NotPetya in 2017 — a wiper disguised as ransomware that spread globally through legitimate accounting-software updates and caused an estimated $10bn in damage, making it the costliest cyberattack in history. Through the full-scale invasion of Ukraine the group has continued deploying purpose-built wipers (SwiftSlicer, ZeroLot), and in August 2023 was linked to Infamous Chisel, malware targeting Android devices used by Ukrainian military units. In January 2026 Microsoft and Amazon reported Sandworm targeting Western critical infrastructure through misconfigured edge devices and VPN appliances. AI use: essentially absent. Wiper malware does not require natural-language fluency; what it requires is operational targeting intelligence and deployment timing, which Sandworm manages through human intelligence networks embedded in Ukrainian institutions.
APT41 — Double Dragon (China / MSS contractors; also Wicked Panda, Winnti, Brass Typhoon; indicted by US DOJ in 2020)
The most unusual major APT: a group running state-sponsored espionage and financially motivated cybercrime simultaneously, apparently with the tacit blessing of China’s Ministry of State Security. Mandiant described the arrangement bluntly — “spies by day, thieves by night” — noting that APT41 leverages classified espionage tooling in what appear to be personal financial operations. The group pioneered supply-chain compromise at scale: its ShadowPad implant, injected into legitimate software updates, became a template used by multiple other Chinese groups. US prosecutors identified links to a Chengdu company called Chengdu 404 Network Technology; five Chinese nationals were charged, and remain at large in China. AI use: limited confirmed evidence. The group’s strength lies in custom malware and pre-positioned access, neither of which benefits much from LLM assistance.
APT35 — Charming Kitten (Iran / IRGC; also Mint Sandstorm, Phosphorus, Magic Hound; active since 2014)
Iran’s premier social-engineering unit, affiliated with the Islamic Revolutionary Guard Corps, spent years cultivating elaborate fake personas — fake journalists, academic researchers, conference organisers — before delivering phishing payloads. In September 2025 internal documents were leaked to GitHub by an unknown party, revealing, for the first time, the bureaucratic machinery behind the charm: operators subject to daily quotas on phishing emails sent, successful compromises achieved and hours spent on reconnaissance, with supervisors filing monthly performance reports that included “phishing success rate” as a formal KPI. The leak also documented a mid-2025 campaign against Israeli journalists and cybersecurity professionals, using AI-polished messages — grammatically flawless, contextually tailored — delivered via WhatsApp to steal credentials to Google accounts. AI use: documented. The group uses AI primarily to eliminate the linguistic and cultural inconsistencies that once made its lures detectable.
Lazarus — TraderTraitor (North Korea / Reconnaissance General Bureau; also APT38, BlueNorOff, Diamond Sleet)
No other threat group has stolen at the scale Lazarus has in recent years. The Bybit heist in February 2025 — $1.5bn in Ethereum redirected through a corrupted signing interface after months of preparation and a social-engineering attack on a Safe{Wallet} developer — was the largest crypto theft on record. The TraderTraitor sub-cluster responsible is characterised by extreme patience: it typically spends weeks or months cultivating targets on LinkedIn under fake recruiter personas before delivering malware-laced job-offer documents. One analyst put the dynamic bluntly: “Lazarus Group is not state sponsored in the traditional sense. Lazarus Group is North Korea.” Estimated DPRK cyber workforce: 8,400 people in 2024, up from 6,800 in 2022. The separate fake-IT-worker scheme — where AI-altered photographs and deepfake video help operatives clear Western job interviews using stolen American identities — runs in parallel, generating tens of millions more in salary fraud. AI use: limited in core crypto-theft operations; confirmed in the identity-fraud layer.
Volt Typhoon (China / PLA-linked; also Bronze Silhouette; targeting US power, water, telecoms, ports)
The contrast case. CISA, the NSA and the FBI issued a joint advisory in 2024 describing an adversary pre-positioned inside American critical-infrastructure networks — in some environments for more than five years — using a tradecraft that reads like a deliberate anti-AI posture: valid stolen credentials, native operating-system tools, no custom malware, no AI-generated reconnaissance, no artefacts of any kind that a model might write. The apparent purpose is not exploitation now but preparation for a future crisis — the ability to degrade power, water or communications at a moment of geopolitical tension. AI use: none confirmed. If anything, AI tooling would raise the forensic footprint and undermine the deniability that makes the operation strategically valuable. Not every frightening thing a state does with a keyboard is an AI story.
Thanks for reading. I added a paid subscription option to this blog. More people becoming paid subscribers will show me that you care, and guilt-hack me into writing on a regular basis, not just when I have the time and inspiration. Thoughtful comments and discussion are an appreciated incentive as well!




Thanks for sharing! I wonder what's your concern for applying for the verification program so that you can use the frontier models without safeguards?
> Restricting frontier AI cybersecurity capabilities continues to be the wrong approach
I don't quite agree with the above point. IIRC, you advocate that defenders should "easily access all of Claude Fable and GPT-5.6’s capabilities as fast as possible", to me that means restricting the capabilities smartly rather than not restricting the capabilities. We still need restriction in the first place otherwise attackers will take lots of advantage. Does this make sense to you?