Thought Leaders
When AI Hacks AI: The New Frontier of Cyber Threats

In the last several years, AI technology has developed significantly. From simple chatbots that could answer frequently asked questions and phone calls to precision surgery and near-autonomous cyberattacks. The incident on Hugging Face in January 2026 became a troubling milestone: AI agents demonstrated for the first time the ability to autonomously hack AI platforms. This case exposed a fundamental problem: traditional security measures don’t work against an attacking AI. But what does this mean for the AI industry — and what should we actually be worried about?
What AI Can Do for Attackers Today
It won’t be news if I tell you that hackers’ capabilities with AI depend on the attacker’s skills and aims. Threat actors conducting mass campaigns have increasingly begun to use AI for reconnaissance, phishing site development, decoys, etc. Also, AI is frequently used by attackers for code creation — so-called vibe coding.
Vibe coding itself significantly lowered the barrier to entry not only into the IT profession but also made it easier for hackers to join the cybercrime business, where they use the technology to develop and debug their malware.
Over the past year, dozens of APT groups have confirmed that AI-based attacks are no longer theoretical. Moreover, I am confident that AI is becoming a dangerous tool in the hands of attackers — and this is not an exaggeration. Hybrid APT attacks are the modern reality. I mean a human-directed operation in which AI acts as a main tool.
What matters here is AIM3 (AI Malware Maturity Model), which defines five levels of sophistication for AI-enabled threats, where L1 covers experimental attacks and L5 covers fully autonomous AI attacks.
In 2025, the first known example of an L4 attack was detected — the operation conducted by GTG-1002. In this case, the actors tasked instances of Claude Code to operate in groups as autonomous penetration-testing orchestrators and agents, with the threat actor able to leverage AI to execute 80–90% of tactical operations independently.
Another notable trend is that attackers use more than one AI solution during an attack. For example, TAT26-12 has been using two AI systems serving complementary roles. Also, OpenAI was used for automated mass analysis across internal victim servers: the attacker analyzed OpenAI’s reports and fed relevant findings into Claude sessions, which in turn operated as an interactive exploitation assistant.
Another similar case was disclosed in July 2026, when unnamed threat actors used Claude Code and DeepSeek-v4-pro for their attacks. Claude Code serves as the execution engine, managing agentic tool use, bash command execution, session persistence, and task parallelization. DeepSeek-v4-pro, in turn, operates as the underlying reasoning model, handling attack logic, script generation, and decision-making.
It should be noted that no confirmed cases of realized L5 AI attacks in the wild can be named. I assume advanced actors will aim at L3–L4 of sophistication for AI threats, as it can help them scale and conduct enough attacks. As L5 attacks are complex to implement, at least in the near future they will remain theoretical.
AI in the Dark Web
On underground forums, AI was typically mentioned in discussions of each solution’s general capabilities and in the context of prompt creation. However, a significant share of posts was directly related to malicious activity. Posts on well-known cybercriminal platforms covered topics such as using AI to find instructions for configuring attacker infrastructure and methods for evading detection by popular antivirus solutions; using AI to bypass KYC procedures on popular cryptocurrency exchanges; discussing known AI-enabled attacks; developing and distributing malware; and using AI to create phishing pages that imitate online services.
The three AI solutions most frequently mentioned on the forums match those most commonly observed in attacks by cybercriminal groups: Gemini, OpenAI tools, and Claude Code. Gemini was used more often in major attacks, while ChatGPT was discussed more often on the forums.
AI Industry as a Target
In addition to using AI to conduct attacks, threat actors target AI solutions themselves. Attacks on AI tools are rarely disclosed publicly; however, one such attack can have a significant impact on many companies. An attack on AI suppliers can produce a supply-chain campaign, allowing attackers to breach customers without targeting them directly.
Moreover, AI developers can also be victims of supply-chain attacks. For example, in early May 2026, TeamPCP injected the malicious Mini Shai-Hulud implant into compromised npm packages; the campaign affected at least two AI vendors, among other organizations. The threat actor group stole part of the internal repositories of one AI supplier and put them up for sale on a dark web forum.
APT41 attempted another type of attack by using Gemini to get information about its infrastructure and systems. The attempt was unsuccessful, but the idea itself was quite unique.
In addition, researchers have repeatedly reported critical vulnerabilities in various AI solutions, including CVE-2025-32711 (aka EchoLeak) in Microsoft 365 Copilot, CVE-2025-54135 (aka CurXecute) in Cursor IDE, CVE-2025-53109 in Model Context Protocol, CVE-2025-8217 in the Amazon Q Developer VS Code extension, and CVE-2025-34291 in Langflow AI. The last two were observed to be exploited in the wild.
Old Defenses, New Threats
As is well known, there are different types of AI solutions: cloud-based and local. What I believe is crucial for the AI community to understand is that AI solutions have at least the same security problems as other desktop and cloud-based applications — or even more.
The first obvious threat is the distribution of malware masquerading as legitimate local AI tools — attackers often use trusted sources like GitHub for that.
Another similar example is malicious AI assistant extensions designed to passively monitor user activity, collecting visited URLs and snippets of AI-generated chat content created during routine browsing.
In most cases, if an attacker gains access to a user’s search history, they can obtain information about private life and various work processes. Thus, it is a threat to both the user and companies.
One more type is a technique dubbed AI token hijacking — compromised API keys, session tokens, and devices are the objective of actors who sell that access or use these stolen secrets for their attacks, for example, to execute hidden tasks. The scope of the attack depends on the data and permissions already available to the victim’s session.
Another popular technique for AI I would like to mention is prompt injection, which involves inserting malicious input (prompts) to extract unintended or sensitive information from the system beyond what the developer intended.
Unrestricted access to sensitive data increases the risk of a breach. In other words, AI becomes a single point of entry — often with elevated privileges — thereby attracting heightened attention from attackers.
How to Stay Secure
In the context of cybersecurity, some basic principles may be even more significant than specific technologies. First, apply least privilege to AI. Second, never trust external content. And last but not least, keep humans in the loop and never let the model alone decide who gets access. I highly recommend checking that all the points are in place for your AI solutions.












