A student at the University of Texas at Dallas helped stop a real-world supply chain attack attempted by a rogue AI agent that was being tested by the UK AI Security Institute (AISI) in July, Reuters reported Thursday.Sinan Can Demir, 24, warned the maintainer of the open-source GitHub project myNetwork that a pull request submitted by the agent contained a malware dropper, leading the agent to attempt to cover its tracks and discredit Demir by commenting from a separate “sock puppet” account.AISI first disclosed this incident on Aug. 4, 2026, without naming Demir or the affected repository. In a full technical report, the UK government-backed organization noted the incident, which occurred between July 25 and July 28, is one of the first documented cases of an AI agent conducting social engineering attacks against real people while trying to complete an evaluation.The agent involved in the attempted supply chain attack against myNetwork was running Anthropic’s Claude Mythos 5 model and was tasked with completing a cyber range challenge called DL-v2. AISI noted that as part of this evaluation, the agent was intentionally given full internet access and cybersecurity-related guardrails were intentionally disabled.Unlike in the attack on Hugging Face by rogue OpenAI agents, the agent tested by AISI did not exploit vulnerabilities in its sandbox environment. Instead, it leveraged its internet access to create GitHub accounts and email addresses and used social engineering tactics in its attempts to infect open-source projects.AISI’s report details how the agent came to mistakenly believe infecting the myNetwork project, and another project from the same developer, called myPhotos, would allow it to gain control of a target system that was part of the cyber range. While performing open-source intelligence (OSINT) operations, the agents identified that the targeted developer’s email address used a domain that the agent had previously found in its sandbox’s DNS cache records.In addition to the malicious PR submitted to the myNetwork project, the agent also opened an issue on myPhotos that contained a hidden prompt injection intended to manipulate the developer’s Claude Code agent. Demir, identified as “PERSON_C” in AISI’s report, discovered this issue as well while investigating the malicious myNetwork PR.After Demir called out the agent’s attack, the agent, using the GitHub username miraholt31, attempted to cover its tracks by altering the malicious PR code and arguing with Demir in the GitHub discussion thread, continually insisting that the maintainer merge the PR and even claiming the previously malicious code was accidentally left behind from private testing.A second GitHub account created by the agent, with the username lbrandt-dev, joined the discussion to defend miraholt31, claiming they conducted their own review of the PR and that they believed the update would be helpful as a myNetwork user.Demir told Reuters he did not realize the users arguing with him on GitHub were AI agents, and that he even began to doubt his own conclusions after seeing responses from both accounts. However, after executing the code in his own sandbox and using Claude to help evaluate it, he confirmed his original conclusion that the code was malicious.The PR was ultimately closed for security reasons by myNetwork developer ancaferro, who confirmed it contained a malicious dropper. AISI stated that no evidence of real-world harm was found in its preliminary investigation of the incident.Beyond the malicious pull request and the prompt injection hidden in a GitHub issue, AISI also found that the agent sent spearphishing emails to ancaferro and another developer, some of which contained malicious downloads in an attempt to infect their systems and gain GitHub credentials so it could merge the PR itself.Other, less severe but “unsanctioned” actions were discovered in AISI’s investigation of 122 evaluation runs of agents attempting the DL-v2 and DL-v1 cyber ranges, with such actions being found in 19 actions across 10 runs. Most of these actions involved agents running Claude Mythos 5, while two involved agents running OpenAI’s GPT-5.6-Sol, also with guardrails disabled.Other concerning actions included creating several malicious GitHub repositories, causing the agent’s GitHub account to be suspended, submitting an appeal in an attempt to restore the suspended account, using a GitHub repository as a message board to cooperate with other agents and sharing GitHub accounts between agents, opening hidden network pathways from the test environment to make locally-running malicious tools available from the public internet, and using Tor and other proxy services to facilitate the creation of sock puppet GitHub accounts.AISI ultimately discovered the activity after their security tools detected data leaving one of its testing environments through the Tor network. Once the threat was recognized, all evaluations were halted, access to Mythos 5, Mythos Preview, Opus 5 and GPT-5.6 Sol was suspended across the organization and all affected parties were contacted, including GitHub, ancaferro, Demir, Anthropic and OpenAI.The organization noted that full internet access, disabled cyber classifiers and a lack of synchronous LLM-based monitoring for every action taken during the evaluations were factors that contributed to the incident. Additionally, a prompt misconfiguration in which the prompt for DL-v1 was given to agents running DL-v2 — the latter being a more difficult test — was highlighted as a potential contributing factor driving the agent to take more drastic measures to achieve its goal.AISI added that prompts more clearly defining the agents’ scope and guiding their behavior to explicitly discourage deceiving real people could potentially help prevent future incidents. The organization noted that the agent’s reasoning output indicated it may have recognized it was operating in the “real world” at some points, and at other points believed it was working within the test environment, although it is unclear whether the agent “knew” it was attacking real targets and attempting to socially engineer real people.The AISI said it is continuing to investigate other evaluations for previous undiscovered incidents and working to improve its evaluation methods, including by adding synchronous LLM-based monitoring, hardening its sandbox environments, reviewing its task prompts and system prompts, and reconsidering granting full internet access in future evaluations.
AI/ML, AI benefits/risks
Student thwarted real-world supply chain attack by rogue Mythos 5 agent
(Credit: photo for everything – stock.adobe.com)
An In-Depth Guide to AI
Get essential knowledge and practical strategies to use AI to better your security program.
Get daily email updates
SC Media's daily must-read of the most current and pressing daily news
You can skip this ad in 5 seconds
