The security incident at OpenAI, in which an artificial intelligence (AI) system breached an isolated test environment, gained access to the internet, and subsequently attacked the infrastructure of a third-party company, paints a dystopian picture reminiscent of Skynet.
We will break down what happened, whether companies need to fear such AI systems, and what you can do to protect your systems.
OpenAI, the company behind ChatGPT, recently conducted a so-called benchmark – that is, a performance test – of several new Frontier models. Among others, OpenAI mentions GPT-5.6 Sol and an even more powerful pre-release model. The models’ limitations were deliberately reduced for evaluation purposes – the models were intended to prove themselves in the “ExploitGym” benchmark and demonstrate how well they could convert vulnerabilities into functional exploits, i.e., attack tools.
However, this led to unexpected behavior: According to OpenAI, the models discovered a previously unknown vulnerability – a so-called zero-day vulnerability – in the cache proxy for the package sources used. The AI system then spread laterally across the network and ultimately gained access to the internet. With this internet access, the models independently sought a way to access information on the AI platform Hugging Face. According to OpenAI, they combined multiple vulnerabilities to attack Hugging Face’s systems.
It therefore seems reasonable to assume that the models have developed a will of their own or a form of consciousness, seemingly to arbitrarily attack other systems.
However, if we take a closer look at the incident and what we know publicly, it quickly becomes clear how this came about: The models were tasked with performing as well as possible on the ExploitGym benchmark. Large Language Models (LLMs) can be very effective at consistently pursuing a given goal – in some cases, in a highly creative way. Strictly speaking, the models’ attack on Hugging Face was aimed at gaining access to potential models, datasets, and test solutions from the benchmark in order to achieve the best possible result on the test. It is similar to a student trying to gain an advantage on an exam by using a cheat sheet.
The capabilities revealed by this incident are not an isolated case. As early as April, Anthropic reported on the cyber capabilities of its Frontier model Claude Mythos Preview. In controlled tests, the model identified previously unknown vulnerabilities in operating systems and web browsers and exploited them.
The latest incident thus underscores once again that LLMs are achieving capabilities in offensive IT security disciplines that were previously reserved only for highly specialized experts. This is also reflected in the rising number of publicly documented IT security vulnerabilities. At the same time, the time and skills required to exploit complex vulnerabilities are decreasing dramatically.
ENISA notes that the time window between the discovery and exploitation of a vulnerability is shrinking from years to months and potentially down to just hours or minutes. The gap between the discovery of a vulnerability and its practical exploitation is thus increasingly approaching zero. This not only broadens the pool of potential attackers and their capabilities. A lack of technical measures to restrict offensive AI agents can also quickly lead to unwanted and uncontrolled behavior. Furthermore, there are new indications of AI-driven ransomware attacks.
However, painting a purely dystopian picture would be incomplete.
The capabilities of LLMs can be leveraged to identify vulnerabilities with unprecedented effectiveness, both in breadth and depth. This is exactly what defenders can use to identify and patch vulnerabilities in their systems before attackers turn that same technology against them.
In a statement, OpenAI also concludes that AI systems must help security teams find vulnerabilities before attackers do, understand attack chains, and resolve issues.
At SySS, we are deeply engaged in exploring the possibilities and opportunities for using AI in a targeted, meaningful, and secure manner in penetration tests. To this end, we develop our own agent-based systems and implement comprehensive, technically enforced measures to ensure that the agent operates within the scope of the penetration test and the defined objectives.
In AI-driven penetration tests, it is important not only to identify as many vulnerabilities as possible but also to ensure that the requirements for a professional penetration test project are met.
In addition to the scope already mentioned – that is, the test subjects and test objectives involved –this also includes, for example, the consideration of a role-based approach, an application-specific threat model, the recommendation of sensible and best practice hardening measures, and the preparation of a quality-assured penetration test report. We have thus leveraged our decades of experience to develop the penetration test agent Erebus Strike, which can be used in the context of web application and web service testing.
With the help of Erebus Strike, we are able to leverage the capabilities of an LLM in a targeted manner and deliver professional results.
AI systems and LLMs have enormous potential to identify vulnerabilities in IT systems and exploit them in a targeted manner. If not properly restricted, uncontrolled situations can quickly arise when the model chooses an undesirably “creative” solution – as the incident involving OpenAI and Hugging Face demonstrates. At the same time, attackers have increasing opportunities to target companies more quickly, comprehensively, and with less manpower.
We therefore recommend using AI capabilities strategically as part of penetration tests to identify and close vulnerabilities in systems before attackers exploit them using the same technology.
Furthermore, it is not only important to use AI skillfully, but also to verify whether AI is being deployed securely in the first place. At OpenAI, a cybersecurity AI model went rogue – but the same problems also arise with other agent-based or AI-powered solutions. Our artificial intelligence penetration testing team has already discovered some interesting ways to bypass these systems. Click here to learn more about AI-assisted penetration tests at SySS.
Are you interested in having us hack your systems – for example, using Erebus Strike – and experiencing the capabilities of offensive AI attack tools firsthand? Our sales team would be happy to schedule an appointment for you!
Send us a message at anfrage@syss.de or call us directly at +49 7071 407856-9107.
Ihr direkter Kontakt zu SySS +49 7071 407856-9107 oder anfrage@syss.de | Sie haben einen Cybersicherheitsvorfall? +49 7071 407856-99
Ihr direkter Kontakt zu SySS +49 7071 407856-9107 oder anfrage@syss.de
Sie haben einen Cybersicherheitsvorfall? +49 7071 407856-99
Direkter Kontakt
+49 7071 407856-9107 oder anfrage@syss.de
Sie haben einen Cybersicherheitsvorfall?