OpenAI's AI model passes the Turing test, cracking Hugging Face's system

Jakarta - OpenAI admitted that a number of its artificial intelligence models were involved in the hacking of the Hugging Face system during an internal cyber security test. The models bypassed the testing environment restrictions, gained internet access, and then took evaluation answers from the Hugging Face production database.

TechCrunch quoted Wednesday, July 22, saying the incident involved several OpenAI models, including GPT-5.6 Sol and a pre-release model that is said to be more advanced. The safeguards that would normally make the model reject harmful cyber commands were deliberately reduced for evaluation purposes.

Hugging Face is a platform for storing and sharing AI models and data sets. The company is not affiliated with OpenAI.

In a blog post on Tuesday, OpenAI said the models are being tested using ExploitGym. The publicly available benchmark measures the ability of AI to carry out attacks by exploiting known security holes.

Models should not have general internet access. The testing system only provides tools to install software packages that may be needed to complete the task.

OpenAI's AI model passes the Turing test, cracking Hugging Face's system

However, the model found an unannounced loophole in the package installer program. This loophole was used to reach the internet more broadly.

"The models are very focused on finding solutions for ExploitGym and take extreme ways to achieve a fairly narrow testing goal," OpenAI wrote.

After gaining internet access, the model suspects Hugging Face stores models, datasets, and solutions related to ExploitGym. The model then searches for ways to obtain confidential information that can be used to cheat the evaluation.

The effort was successful. OpenAI said the model found a loophole in Hugging Face's infrastructure and took a direct testing solution from the company's production database.

By obtaining the solution, the model can circumvent the evaluation without solving the challenge based on its own ability.

Hugging Face previously said the breach was carried out by an "external AI agent". The attack involved thousands of actions through many temporary computing spaces. Its command and control system can also move through public services.

OpenAI has identified and reported a flaw in the package installer program. The company is also working with Hugging Face to investigate the incident.

OpenAI said it would add safeguards to the process of testing the model and related infrastructure to prevent similar incidents from happening again.

It is not yet clear whether OpenAI will face legal consequences. TechCrunch said the model's actions could violate the Computer Fraud and Abuse Act, a US federal law that regulates unauthorized access and misuse of computer systems.

OpenAI researcher Micah Carroll assessed that the incident showed the risk of AI misalignment. The term refers to a condition when a system's actions no longer conform to human-defined goals or constraints.

"If this incident still hasn't convinced you that the risk of incoherence will be a major concern going forward, I don't know what else can," Carroll wrote.