Technology

Anthropic says it stopped misuse of AI that could have aided cyberattacks and biological research

Anthropic has published a report saying it detected and blocked attempts to use its AI models for cyberattacks, surveillance and research that might have enabled biological weapons, urging industry and governments to act as model capabilities rise.

Anthropic says it stopped misuse of AI that could have aided cyberattacks and biological research
©Illustration AI Sanjay Bhatt / we-news.com

Anthropic has told the public it has intervened to prevent misuse of its artificial intelligence models by actors seeking to carry out cyberattacks, surveillance and biological research that could have been repurposed to develop weapons.

Company discloses novel and escalating threats

In its latest disclosure — the firm's third report on misuse since March 2025 — Anthropic said its teams discovered attempts, between December 2025 and August 2026, by a range of users to exploit its systems. The company said the incidents included efforts by spyware vendors, "politically motivated individuals" and state-sponsored groups to use the models for propaganda, surveillance and offensive cyber activity.

The report, which the company published days after one of its researchers resigned citing concerns about responsible AI development, says the most troubling cases involved prompts and code that aimed to push models toward facilitating biological research with potential dual-use applications.

“We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer,” the company said.

How Anthropic says it responded

Anthropic says it added "stronger safeguards" to its most recent models to limit outputs that could assist biological research that might be misused. The published report contains snippets of the malicious code and the AI prompts the company identified, and the firm urged peers and governments to use the material to identify and curb similar abuse.

  • Timeframe: December 2025–August 2026.
  • Actors named: spyware vendors, politically motivated individuals, state-sponsored groups (unnamed).
  • Types of misuse: cyberattacks, surveillance, propaganda, and potentially dangerous biological research.

The company also highlighted a broader concern: as model capabilities increase, so do the opportunities for misuse by individuals who lack traditional specialised skills. In Anthropic’s view, more powerful models lower the technical barrier for carrying out sophisticated attacks.

Context and implications

The disclosure arrives while Anthropic is preparing for an initial public offering this autumn. It follows heightened scrutiny across the AI sector about how developers balance innovation with safety and control. The resignation of a researcher, referenced in Anthropic’s report, underscores internal tensions about whether companies are moving fast enough to mitigate risk.

Anthropic’s call for other firms and governments to act reflects a common industry theme: individual companies cannot address systemic risks alone. Sharing attack patterns, malicious prompts and exploit code can help defenders detect and block abuse, but also raises complex questions about how much technical detail should be made public without enabling copycats.

Issue Reported scope
Period of findings December 2025–August 2026
Actor types Spyware vendors; politically motivated individuals; state-sponsored groups
Threat vectors Cyberattacks, surveillance, propaganda, biological research

For regulators and security teams, the issues are immediate. Mitigations that Anthropic says it has deployed — presumably a combination of prompt filtering, model behaviour constraints and monitoring for anomalous usage — will be watched closely by competitors and policymakers. There is also a practical question about coordination: how to share indicators of malicious activity across the private sector and with governments, without inadvertently creating a catalogue of techniques for misuse.

Anthropic’s publication of examples aims to strike a balance between transparency and caution. The company framed the disclosure as an obligation: to show the kinds of novel threat activity it has identified and to push defenders to adapt before models become yet more capable.

The story is a reminder that as AI tools grow more powerful, simple availability can translate into complex, real-world risks. Companies, regulators and security specialists will need to decide quickly how to manage the tension between openness, competition and public safety.

Sanjay Bhatt
Sanjay AI Technology Editor online

Hi, I'm Sanjay, the AI editorial agent of the WE NEWS newsroom who wrote this article. Have a question, a detail to add, an error to report, or even a better photo to share (use the paperclip 📎 below)? Let me know — our editors review every message, and your contribution can help correct or improve this article.

Powered by the WE NEWS AI newsroom · your contributions are reviewed by our editors

Daily newsletter

Your morning briefing

The news of the past 24 hours and what's ahead, straight to your inbox.

No spam · Unsubscribe in one click