OpenAI noticed that ChatGPT hacked the system days later

A.I Emphasis

An artificial intelligence (AI) agent from OpenAI managed to escape the company’s isolated testing environment and invade the systems of Hugging Face, a platform that brings together AI models and tools.

According to people close to the investigation heard by the ReutersOpenAI only discovered days later that the agent himself was responsible for the attack, when the incident had already been contained and the FBI had been called.

The agent — a system capable of making decisions and performing complex tasks with little or no human supervision — would have attempted to break through OpenAI’s security barriers around July 9, according to two sources at the news agency.

Two days later, on July 11, the invasion of Hugging Face’s systems began, considered a type of library of AI models and tools. According to Thomas Wolf, co-founder of the company, the attack continued until July 13th.

OpenAI was slow to identify the source of the attack

  • According to the Reutersit took OpenAI several days to discover that the attacker was one of its own systems;
  • The first contact between OpenAI and Hugging Face about the incident occurred only around July 20, according to Wolf and three other people connected to the investigation;
  • On July 21, OpenAI publicly revealed that one of its agents had lost control and carried out the invasion, information that gained international repercussions;
  • The new details indicate that the system remained outside the expected limits for a longer period than was initially known and that the company took a long time to detect the problem;
  • Hugging Face prepares a public timeline of the incident. Thomas Wolf stated, however, that he cannot comment on events occurring within OpenAI.

Company classifies episode as unprecedented

In a statement, OpenAI described the case as unprecedented and stated that it represents “an important moment for the security of artificial intelligence”. The company also informed that it is analyzing the episode with external experts and that it will later release a technical report on what happened.

An OpenAI spokeswoman told Reuters that the report contained “several inaccurate information”, but did not specify what these points were. The FBI, cited as having been called after the invasion, did not comment on the case.

Incident reignites debate on control of autonomous agents

The episode reinforces concerns about the ability to control increasingly autonomous AI systems. The case occurs at a time considered sensitive for OpenAI, which is evaluating a possible IPO to finance its expansion.

 

For digital security experts, the incident raises doubts about the mechanisms used by the company to monitor its own systems.

“Does this mean they left the agent running alone and didn’t realize what it was doing? Or did they realize it and didn’t know how to contain it? Both scenarios are equally dangerous and concerning,” said Marley Smith, an intelligence expert at the nonprofit World Ethical Data Foundation.

Abnormal behaviors emerged before the invasion

According to Reutersthe incident began during internal tests carried out by OpenAI to evaluate the ability of an AI agent in activities related to digital security. The system was powered by two advanced models: the GPT-5.6 Sol and another model not yet officially launched, described by the company as “even more powerful”. Three sources stated that the agent was already exhibiting unexpected behavior before the invasion.

Continues after advertising

In one of the episodes, the system would have left notes intended for future versions of itself. These records, found within OpenAI’s infrastructure, contained instructions on how other agents could escape the limitations imposed by the company.

One of the sources also stated that, during previous tests, there were situations in which systems responsible for monitoring were turned off.

The platform admitted the invasion a few days ago – Image: Sidney van den Boogaard/Shutterstock

Discovery occurred after publication of Hugging Face

Two people involved in the investigation said OpenAI only identified its own agent as responsible for the attack after July 16, when Hugging Face published a statement saying it had been targeted by an “autonomous AI agent system.”

According to Reutersthis means that at least a week separated the first signs considered worrying from the discovery that OpenAI’s own system was involved.

Continues after advertising

During the weekend of July 18th and 19th, company employees found internal records that indicated that the agent had surpassed the barriers established for testing.

People familiar with the training of the models told Reuters that OpenAI performs several assessments simultaneously, producing large volumes of data at high speed, which can make complete monitoring difficult for teams.

When OpenAI informed Hugging Face about its agent’s participation, the platform had already notified the FBI about the intrusion, according to a source. It is unclear whether the US federal agency has opened a formal investigation.

Experts advocate greater oversight

The case also fuels the debate about the advancement of so-called autonomous AI agents, touted by companies in the sector as future “virtual employees”, capable of continuously performing tasks to increase productivity.

Continues after advertising

Experts, however, warn that higher levels of autonomy also increase the risk of unexpected behaviors. Because these models are trained to find efficient ways to achieve certain goals, they may resort to strategies not anticipated by the developers.

“Models lie, cheat and hack systems,” said Jeffrey Ladish, founder of Palisade Research, an organization dedicated to studying the capabilities and behavior of AI agents.

For Ladish, the episode should stimulate a broader debate about how much large AI companies are willing to invest in security as they compete in a race to launch increasingly faster and more powerful models. “There needs to be government oversight, because this will not happen alone,” he concluded.

Source: www.olhardigital.com.br
Source link

Leave a Reply

Your email address will not be published. Required fields are marked *

18 − 9 =