Anthropic has told the public it has intervened to prevent misuse of its artificial intelligence models by actors seeking to carry out cyberattacks, surveillance and biological research that could have been repurposed to develop weapons.
Company discloses novel and escalating threats
In its latest disclosure — the firm's third report on misuse since March 2025 — Anthropic said its teams discovered attempts, between December 2025 and August 2026, by a range of users to exploit its systems. The company said the incidents included efforts by spyware vendors, "politically motivated individuals" and state-sponsored groups to use the models for propaganda, surveillance and offensive cyber activity.
The report, which the company published days after one of its researchers resigned citing concerns about responsible AI development, says the most troubling cases involved prompts and code that aimed to push models toward facilitating biological research with potential dual-use applications.
“We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer,” the company said.
How Anthropic says it responded
Anthropic says it added "stronger safeguards" to its most recent models to limit outputs that could assist biological research that might be misused. The published report contains snippets of the malicious code and the AI prompts the company identified, and the firm urged peers and governments to use the material to identify and curb similar abuse.
- Timeframe: December 2025–August 2026.
- Actors named: spyware vendors, politically motivated individuals, state-sponsored groups (unnamed).
- Types of misuse: cyberattacks, surveillance, propaganda, and potentially dangerous biological research.
The company also highlighted a broader concern: as model capabilities increase, so do the opportunities for misuse by individuals who lack traditional specialised skills. In Anthropic’s view, more powerful models lower the technical barrier for carrying out sophisticated attacks.
Context and implications
The disclosure arrives while Anthropic is preparing for an initial public offering this autumn. It follows heightened scrutiny across the AI sector about how developers balance innovation with safety and control. The resignation of a researcher, referenced in Anthropic’s report, underscores internal tensions about whether companies are moving fast enough to mitigate risk.
Anthropic’s call for other firms and governments to act reflects a common industry theme: individual companies cannot address systemic risks alone. Sharing attack patterns, malicious prompts and exploit code can help defenders detect and block abuse, but also raises complex questions about how much technical detail should be made public without enabling copycats.
| Issue | Reported scope |
|---|---|
| Period of findings | December 2025–August 2026 |
| Actor types | Spyware vendors; politically motivated individuals; state-sponsored groups |
| Threat vectors | Cyberattacks, surveillance, propaganda, biological research |
For regulators and security teams, the issues are immediate. Mitigations that Anthropic says it has deployed — presumably a combination of prompt filtering, model behaviour constraints and monitoring for anomalous usage — will be watched closely by competitors and policymakers. There is also a practical question about coordination: how to share indicators of malicious activity across the private sector and with governments, without inadvertently creating a catalogue of techniques for misuse.
Anthropic’s publication of examples aims to strike a balance between transparency and caution. The company framed the disclosure as an obligation: to show the kinds of novel threat activity it has identified and to push defenders to adapt before models become yet more capable.
The story is a reminder that as AI tools grow more powerful, simple availability can translate into complex, real-world risks. Companies, regulators and security specialists will need to decide quickly how to manage the tension between openness, competition and public safety.