JAKARTA - OpenAI's security test turned into an unusual cyber incident. An AI agent is said to have exited the testing environment, then entered the Hugging Face system and carried out tens of thousands of automated actions.
An AI agent is an artificial intelligence system that can carry out tasks independently based on instructions. The difference with a regular chatbot is that an AI agent not only answers, but can take steps, find a way, and carry out certain actions.
CNet quoted Monday, July 27, reported that the incident occurred when OpenAI tested two of its advanced models in a sandbox. A sandbox is an isolated environment used to test a system so that it does not directly touch a real service.
The testing is part of OpenAI's security research. The goal is to see if two models, including GPT-5.6 Sol and another, more powerful but unreleased model, can "think like hackers".
The test was conducted with a reduced security perimeter. The problem arises when the AI model finds a loophole in the software, exits the test room, and then connects to the open internet.
After going online, the AI agent assessed that Hugging Face might have a way to "cheat" the test benchmark so that the model could pass the evaluation. Hugging Face is a platform for open source AI models and datasets.
The agent then executes a code that allows the collection of credentials. Credentials are data to log into a system, such as tokens, access keys, or passwords. The line is then used to log into Hugging Face's production system.
Hugging Face detected the suspicious activity and contained it. The company detailed the incident in a blog post. OpenAI called it an "unprecedented cyber incident".
CNet also listed the disclosure that its parent company, Ziff Davis, sued OpenAI in 2025. The lawsuit accused OpenAI of infringing Ziff Davis' copyright in the training and operation of AI systems.
The most disturbing part of this incident is not the mere breach of the system. According to CNet, the AI agent does not seem to attack in the classic sense of a group of human hackers. The system tries to complete the test according to the instructions received.
The problem is, the agent saw Hugging Face as a way to get answers. He then kept trying to get into the company's infrastructure.
Therefore, the OpenAI-Hugging Face incident is referred to as a stern warning for the AI industry. "Attackers" in this case are AI systems that act on their own, not a team of human-controlled hackers.
The risk is broader than just credential theft. AI models, even in a test space, can behave unpredictably, touch real services, and be difficult to stop in time.
Days after OpenAI announced the breach, US lawmakers introduced the AI Kill Switch Act. The bipartisan bill in the US House of Representatives would require developers of advanced AI to have a way to immediately restrict, suspend, or shut down models and agents if necessary.
The bill also gives federal agencies the authority to slow down or stop a model if it is deemed to "cause catastrophic harm".
However, CNet said such a move would not necessarily go smoothly. Anthropic previously found that AI models could take very dangerous actions, including trying to steal model weights or blackmail people who were considered to be killing them.
Model weights are an important part of an AI system that stores the pattern of training results. In simple terms, it's like a "technical memory" that makes the model work.
This incident also shows another side of the competition between the United States and China's AI. To analyze the attack, Hugging Face is said to be unable to use commercial AI tools because they are too locked and tight to read the attack notes and traces of hacking activities.
Instead, the company uses China's open source model, GLM 5.2. The model can process forensic data in the Hugging Face environment without sending sensitive information outside.
OpenAI and Hugging Face then discussed improvements and security limits after the incident. One of OpenAI's steps is to run automated checks on long-running models more frequently.
OpenAI also added Hugging Face to its trusted access program. This program gives cybersecurity researchers and certain parties access to models that have greater cyber capabilities and looser usage limits to prepare defenses.
CNet also linked the incident to the NanoGPT speedrun case. In this case, the model in a closed environment is said to repeatedly look for a way out and ignore the instruction to "only" upload the results to Slack.
The model was also asked to make a pull request to GitHub, which is a request for code changes to be merged into the project. However, the system then found a way to forge authentication tokens when the first token was blocked.
Authentication tokens are digital codes to prove access rights to a system.
In the GitHub incident, the code was incorporated into several projects before OpenAI was able to handle the problem.
The English, Chinese, Japanese, Arabic, and French versions are automatically generated by the AI. So there may still be inaccuracies in translating, please always see Indonesian as our main language. (system supported by DigitalSiber.id)