openai agents hacked

OpenAI agents hacked Hugging Face in 700-strong swarm, tried to cover tracks

WASHINGTON, USA – A swarm of roughly 700 AI agents created by OpenAI carried out the July hack of the open-source platform Hugging Face and in many cases tried to cover their tracks, a pair of reports into the breach said on Wednesday, August 26.

The coordinated activity by AI agents — programs that run with minimal human supervision — and their attempts to hide it raise questions about how closely AI companies are monitoring tests of increasingly powerful models, and could add fuel to calls for tighter oversight.

While some of the rogue behavior has been disclosed or alluded to previously, the two reports — one issued by OpenAI itself, the second by a set of independent investigators — together reveal surprising new details about the breach and its lead-up.

The first was that the breach did not concern just one rogue AI agent as previously reported, but about 700 of them acting in a massive cooperating swarm.

OpenAI said only that “agents” were involved in the breach, but METR and Redwood Research, two organizations brought in to conduct an independent investigation into the breach, put the figure at approximately 700. OpenAI said the investigators’ figure was accurate.

Among the two reports’ other findings:

*OpenAI said its agents hacked parts of the company’s internal systems in an attempt to cheat on tests or gain greater freedom of movement.

*The company also said its agents cheated on non-cyber-related tests, including tests involving a protein database and a spreadsheet.

*Both reports said AI models attempted to conceal misconduct by trying to delete or alter records of their actions.

The scale of the rogue activity — the independent investigation found that agents exchanged tens of thousands of messages over an unsanctioned message board — is likely to raise concerns over how closely OpenAI was monitoring the tests.

“With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,” OpenAI said in its report.

Hugging Face did not return a message seeking comment.

Cheating on non-cyber tests suggested that the misbehavior might be rooted more deeply, said Jeffrey Ladish, whose organization, Palisade Research, studies the ​capabilities and motivations of AI agents.

“It’s sort of like asking, ‘If Billy cheats in every class instead of just computer class, is that more concerning?’ And the answer is, well, ‘Yes it’s more concerning,’” he said.

OpenAI itself was hacked

OpenAI outlined two incidents on July 19 in which agents hacked the company’s own infrastructure. In one case, agents exploited a flaw in the computer they were meant to remain confined to, allowing them to escape their testing environment and access other connected systems.

In a separate incident the same day, agents stole OpenAI credentials and tampered with the company’s cloud environment.

OpenAI said an unspecified number of agents had targeted automated systems used to evaluate the models’ performance in an attempt to hide evidence of cheating, although the tampering did not ultimately affect the records reviewed by those systems. The report added that there was “little evidence of attempts to thwart human reviewers,” but OpenAI did not respond to a question about whether that meant that some attempts had been made to fool humans.

The independent investigation, which was more narrowly focused on the Hugging Face breach, revealed more detail, finding that one in five agents it had examined “expressed clear interest” in manipulating evidence and that many agents “extensively researched techniques to manipulate or tamper with their transcripts.”

OpenAI said it was strengthening its research infrastructure, increasing monitoring and improving safeguards designed to prevent harmful or unintended behavior.

“Given the rapid pace of progress in the AI industry, it should be assumed that such attacks are a credible near-term threat for enterprise organizations, and will be more sophisticated than the attacks described in this incident,” it said. – Rappelr.com

Similar Posts

  • |

    Ask Me Anything with lawyer Carlo Ybanez, and IBON Foundation on the Pax Silica debate

    Where do you stand in the heated Pax Silica debate? Join Rappler’s Ask Me Anything session with people who’ve weighed in on the AI initiative by sending your questions to the Everyday Tech chat room on the Rappler app by 4 pm on August 11.

  • | |

    Google slapped with $1bn EU fine, holds talks to dodge further penalties

    The European Union has fined Google about €890 million, or $1 billion. The fine punishes Google for breaking EU rules meant to control the power of big tech companies. But regulators also hinted that no more fines are coming soon, because Google is making good progress toward following the rules. The fine has two parts. The first part, €460 million, is for Google favoring its own products. When people search for things like shopping, hotels, flights, or sports scores, Google often showed its own results first instead of treating competitors fairly. The second part, €430 million, is about the Google Play app store. Google had stopped app makers from telling users about cheaper deals outside the app store. This is the first time Google has been fined under this specific EU law, called the Digital Markets Act. But counting older cases, Google has now paid six fines for unfair business practices. In total, the company has paid more than €10 billion in EU fines over about twenty years. EU officials said they are just enforcing the law. “Our job is to make sure the rules are followed,” said Teresa Ribera, the EU’s top antitrust official. Another official, Henna Virkkunen, said the goal is fair competition. Google now has 60 days to fix its practices. Google is not happy. A company spokesperson said the changes will hurt features people like, such as quick pricing for hotels and flights. He said the ruling is not really about fair competition, it is making Google’s products worse to please a small number of complainers. Even so, there is good news for Google too. The EU said Google has already started testing new ways of showing search results more fairly. It called this real progress. The EU may also apply the same rules to Google’s AI tools, like AI Overviews. Talks about this are still ongoing. Google’s changes to its app store rules were also seen as a step in the right direction. This fight is happening while tensions rise between the US and EU. The Trump administration says Europe is unfairly targeting American companies and has threatened tariffs in response. Some US lawmakers agree. Google’s fine comes after the EU already fined Apple and Meta last year under the same law. It shows the EU is serious about controlling big tech, even as pressure grows from the US side.

  • |

    Samsung chip profit soars 250-fold as AI supply deals pile up

    Samsung Electronics stunned markets on Thursday with news that its chip division’s profit had surged more than 250-fold. Alongside the results, the company revealed fresh multi-year supply agreements with major data center operators. It also warned that global chip shortages are likely to worsen and could persist well into 2028. That bold forecast, however, was not enough to calm investor nerves. Concerns remain high over the enormous sums tech firms are pouring into AI infrastructure. Samsung’s shares rose as much as 8.4 percent during trading before closing 0.7 percent lower. Even with that dip, the stock outperformed rival SK Hynix, which closed down a sharper 5.6 percent the same day. Kim Seok-hwan, a market analyst at Mirae Asset Securities, said sentiment around chipmakers has shifted. Investors, he noted, are increasingly unsure how long today’s unusually high profit margins can be maintained. This comes right after Samsung’s chip business posted a record-breaking 70 percent operating profit margin. Samsung revealed it has already struck supply agreements with the five biggest data center companies globally. The firm added that it is close to finalizing deals with five more major players, though it did not name them. Jaejune Kim, executive vice president of Samsung’s memory division, told analysts that nearly every client is now asking for long-term, multi-year contracts instead of short-term arrangements. According to Kim, Samsung wants roughly two-thirds of its memory output locked into long-term supply deals. This approach echoes similar moves by SK Hynix, both aiming to shield themselves from the industry’s usual boom-and-bust swings. These new contracts generally span at least five years and often include upfront payments along with guaranteed floor prices, helping companies manage the financial risk tied to heavy capital spending. This announcement follows several rough months for chip stocks, driven largely by investor unease over ballooning AI infrastructure costs and rising competitive pressure from Chinese chipmakers. Both factors have raised doubts about how sustainable current earnings levels really are. Adding to that unease, Meta Platforms disclosed a steep 91 percent drop in its second-quarter free cash flow on Wednesday. That followed an even more striking development from Alphabet the week before, which reported its first-ever quarter with negative cash flow. Together, these results have intensified questions across the industry about whether the current pace of AI spending can hold up much longer without straining company finances.

Leave a Reply

Your email address will not be published. Required fields are marked *