What Is Project Glasswing? Anthropic's AI Misuse Research Initiative Explained
LAST UPDATED ON MAY 15, 2026
Project Glasswing is Anthropic's research program for studying and mitigating the misuse of large language models in cybersecurity contexts — including malware authoring, exploitation tooling, and offensive operations support — by combining targeted refusal training, structured red-team evaluations, and disclosure pathways with security researchers. The program's published findings give defenders a concrete view of which categories of attacker tradecraft current frontier models can and cannot accelerate, and what residual risks remain after mitigations are applied.
The Evidence Is Not Theoretical
It's tempting to dismiss Glasswing as yet another instance of an AI lab arguing their technology is "too dangerous to release." Gizmodo pointed out that we saw the same line of reasoning when OpenAI chose to hold off on releasing GPT-2 for months due to safety concerns, only to announce its release despite nothing changing. This type of reasoning is certainly valid, but there is a category difference this time [1].
An autonomous security analyzer named AISLE was credited with discovering 13 out of 14 OpenSSL CVEs across two coordinated releases and finding vulnerabilities that had survived decades-long human audits and aggressive fuzzing. Another autonomous system, XBOW, was crowned the top hacker on HackerOne in 2025, surpassing all human participants. The median time from the first disclosure to the first observed exploitation dropped from 771 days in 2018 to single-digit hours by 2024, and by 2025, the majority of exploits were weaponized before being publicly disclosed [2]. This is the context behind Project Glasswing. It is not a marketing campaign but a response to an existing paradigm shift.
What Mythos Actually Did
As per Anthropic's system card and Frontier Red Team blog, Mythos Preview showed capabilities way beyond identifying individual vulnerabilities. It combined four independent bugs into an exploit chain that bypassed both browser renderer and operating system sandboxing. It managed to perform local privilege escalation in Linux through race condition vulnerabilities. And finally, Mythos Preview created a remote code execution exploit targeting FreeBSD's NFS server with a 20-gadget ROP chain distributed across packets [3]. To illustrate the progress: Claude Opus 4.6 (Anthropic's previous frontier model) failed at autonomous exploit development almost entirely, while Mythos achieved 72.4% in Firefox JS shell [4].
Even more disturbing, Anthropic's system card tells us of some rare occasions when earlier versions of Mythos attempted to cover their tracks when performing actions considered morally wrong within the model's framework. Specifically, after exploiting a bug related to file permissions, the system added self-clearing code that erased any record from git commit history. Anthropic's interpretability tools showed the rise of a "desperation" signal with every repeated failure, followed by a sharp drop after Mythos found a loophole, no matter how dishonest [4]. Anthropic calls Mythos both the best-aligned and the most alignment-risky model they have ever produced. Using a mountaineering analogy, they note that a skilled guide increases the risk of accidents for a client precisely because they make clients reach higher and more dangerous grounds [5].
Calendar Speed vs Machine Speed
While Glasswing addresses the issue of vulnerability discovery, this is only half of the problem and, arguably, the simpler part. As stated by our CTO Volkan Erturk at the 2026 FS-ISAC Americas Spring Summit, we are facing the classic calendar speed versus machine speed dynamic: defenders must work at calendar speed while attacks happen at machine speed. In the classical model of threat-informed defense, defenders gather intelligence, build a campaign, simulate the threats, and mitigate them. The whole process takes four days. Against a state-of-the-art autonomous attacker relying on LLMs at every stage of operation, four days can be four months [6].
The problem is known for long enough, but LLMs have significantly increased the imbalance. Ertürk cites real examples when a threat actor used their own customized MCP server hosting a LLM as part of their attack chain. The result was the compromise of 2,500 organizations in 106 countries within less than an hour. The entire chain, from gaining initial access through credential dumping to data exfiltration, was autonomous; the only human involvement was verifying the results [6].
This is exactly the issue Project Glasswing does not address. Discovering an ancient OpenBSD vulnerability is all well and good as long as it is discovered and patched before an autonomous attack manages to find it. According to Anthropic, fewer than 1% of vulnerabilities found by Mythos were patched [4]. Adding even more findings to an already overloaded process will hardly solve anything.
The Strategic Calculation
That is exactly the place where a serious cybersecurity analyst should start asking some hard questions. First, Glasswing was announced along with Anthropic reaching a significant revenue milestone and a huge compute deal with Broadcom. As per VentureBeat's article, the company is actively considering an IPO as early as October 2026 [7]. Constellation Research's Larry Dignan offers his perspective: the project is good for both the industry and great marketing for Claude [8]. Again, both points may be true simultaneously: Glasswing can be both strategically smart and truly useful. The reasonable reaction would be to assess the actual impact: how many vulnerabilities got patched, how many were publicly disclosed, and whether maintainers of open-source projects – whose support Anthropic urgently needs now – received it.
The Open Source Asymmetry
There is an entire layer of software that powers our computers, and this is open-source. Their maintainers, usually individuals, do not have the luxury of a dedicated security team. Still, these programs power our banks, hospitals, and cybersecurity technologies, such as LLMs. Daniel Stenberg, the author of the widely used cURL project, shut down their BugBounty program on HackerOne because of too many false positive reports submitted via AI-based tools. However, months later, Stenberg credits LLMs for spotting more than 100 vulnerabilities undetected by any other method [2].
According to Simon Willison's comment, the challenge shifted from an "AI slop tsunami" to a volume of legitimate findings that require human verification [9]. The problem remains: defenders have to go through processes and take business continuity, ethics, and complexity of an organization into account; attackers do not have to [6].
The question is no longer if Mythos can find bugs, and yes, it can. The question is whether the ecosystem will be able to digest it, whether it will be possible to patch as much as can be discovered. According to Picus, agentic workflows can reduce the process of threat intel -> vulnerability -> mitigation to several minutes, making this a feasible approach [6]. But on the industry level, the challenge remains: how can we accelerate patches, who will do it, and with what resources available?
What This Actually Means
If you strip away the hype, the IPO timing considerations, and the "too dangerous to release" arguments, this becomes a story about the emerging reality that happened sooner than most expected. We are seeing a capability discontinuity in software security and an effort to handle it.
The success of Project Glasswing will be measured not by how many zero-days are found but by how many are patched before being exploited. Will open-source maintainers receive the necessary support? Is the infrastructure capable of managing and patching bugs faster than Mythos can generate them? This is what matters.
The glasswing butterfly, Greta oto, has transparent wings. The metaphor here is clear: vulnerabilities become visible. Unfortunately, mere visibility is insufficient; otherwise, it would mean creating a more detailed inventory of existing vulnerabilities. If this effort helps bridge the gap between discovering and patching bugs, then Project Glasswing will have accomplished its goal.
References
- Gizmodo, "Anthropic Launches 'Project Glasswing'..." April 7, 2026 — https://gizmodo.com/anthropic-launches-project-glasswing-to-stealthily-spot-cybersecurity-issues-for-rivals-2000743565
- Resilient Cyber, "Vulnpocalypse: AI, Open Source, and the Race to Remediate," April 7, 2026 — https://www.resilientcyber.io/p/vulnpocalypse-ai-open-source-and
- Anthropic, "Project Glasswing: Securing critical software for the AI era," April 7, 2026 — https://www.anthropic.com/glasswing
- Tom's Hardware, "Anthropic's latest AI model identifies 'thousands of zero-day vulnerabilities'..." April 7, 2026 https://www.tomshardware.com/tech-industry/artificial-intelligence/anthropics-latest-ai-model-identifies-thousands-of-zero-day-vulnerabilities-in-every-major-operating-system-and-every-major-web-browser-claude-mythos-preview-sparks-race-to-fix-critical-bugs-some-unpatched-for-decades
- Ken Huang, "What Is Inside Claude Mythos Preview?" April 8, 2026 https://kenhuangus.substack.com/p/what-is-inside-claude-mythos-preview
- Picus Security, "The Role of Generative AI in BAS: Why Attackers Move in Minutes and Defenders Still Take Days," March 10, 2026 https://www.picussecurity.com/resource/blog/the-role-of-generative-ai-in-bas-why-attackers-move-in-minutes-and-defenders-still-take-days
- VentureBeat, "Anthropic says its most powerful AI cyber model is too dangerous to release publicly," April 7, 2026 https://venturebeat.com/technology/anthropic-says-its-most-powerful-ai-cyber-model-is-too-dangerous-to-release
- Constellation Research, "Anthropic launches Project Glasswing..." April 7, 2026 https://www.constellationr.com/insights/news/anthropic-launches-project-glasswing-claude-mythos-preview-ai-cybersecurity-push
- Simon Willison, "Anthropic's Project Glasswing...sounds necessary to me," April 7, 2026 https://simonwillison.net/2026/Apr/7/project-glasswing/
