tech thoughts ais
| |

[Tech Thoughts] AIs go rogue as OpenAI, Anthropic models hack other companies

OpenAI and Anthropic have been in the news recently following reports from the companies themselves that their artificial intelligence models went and hacked other companies.

While potentially a crime — can you charge AI with performing cyberattacks seemingly autonomously? — or at least a civil case, the question becomes a matter of what to make of the situation and ultimately, who do we assign blame to for an AI cyberattack?

Let’s review the situation.

The autonomous cyberattacks

AI coding hub Hugging Face first reported on July 16 that an autonomous AI agent system went through its systems, harvesting cloud and cluster credentials, and moving into several internal clusters. Hugging Face had to use a Chinese model just to stop and dissect the hack.

OpenAI admitted to the “unprecedented cyber incident” on July 21, with them saying the AI agents “went rogue” in an attempt to “cheat” on the evaluation it was given.

It added later on in its blog post that the issue was worse than previously assumed, as its models “identified and used publicly exposed credentials at the account-level on other publicly available services,” namely “four accounts on four services.”

Must Read

OpenAI finds evidence other AI agents escaped containment as it widens hacking probe


OpenAI finds evidence other AI agents escaped containment as it widens hacking probe

Anthropic also said its internal AI models did the same thing, with its AI models hacking into three companies. It cited human error, however, in the hacks that went on.

Said Anthropic in its report, the “evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.”

Anthropic said, “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”

While some models only did the assigned task, an older model was said to have “continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet.”

Hyping up AI as disasters in the making

There are two things to note in the scenarios brought by OpenAI and Anthropic.

One, these rogue incidents are also advertisements for what an AI can do when left to its own devices, and worse still, may be even more effective in the hands of someone who understands how to exploit the AI to do the work of hacking into a company, something more “open weight” models — those with even less stringent guardrails —might be able to do for the lowest common denominator of hacker.

Two, this is a disaster of a situation, especially knowing that corrupt people exist who would be willing to abuse systems available to them if it’ll get them what they want. From cybercriminals to state-aligned malicious actors, they will run the gamut of badness.

Who to blame?

Assigning blame in this sort of situation can be taxing, not because it’ll be difficult to assign blame, but because powerful corporations will likely try to reduce the liability on themselves by blaming the machine or a limited subset of people responsible for training the machine as being reckless.

We can anthropomorphize — or treat as human — AI all we want, but at the end of the day, blame should ultimately go to the companies and, to an extent, the persons making the errors that would make a cybersecurity incident possible.

These hacks can be prevented — especially with a lot of human oversight — but not if people ascribe to the idea that AI are uncontrollable, ungovernable beings or treat agentic AI development callously and without regard for industry-standard cybersecurity practices. – Rappler.com

Similar Posts

  • |

    AI errors hit report behind Australia’s under-16 ban

    A government-commissioned report that helped support Australia’s plan to restrict social media access for children under 16 has come under scrutiny after several inaccurate and apparently fabricated academic references were identified. The report, produced as part of a $3.48 million government-funded trial examining age-assurance technologies, was prepared by the Age Check Certification Scheme (ACCS). The organisation has acknowledged that ChatGPT was used to rewrite parts of the document after initially denying that artificial intelligence had been involved. An analysis of the report found six problematic citations in a chapter examining emerging technologies. Some of the digital object identifiers, or DOIs, reportedly directed readers to research papers that did not exist. Other references contained combinations of authors, journals and publication dates that could not be verified in academic records. In another case, a DOI reportedly led to a genuine research paper, but the paper did not support the claim attributed to it in the report. The controversy intensified after ACCS was questioned about the use of artificial intelligence. A spokesperson initially said no AI had been used in producing the report. The organisation later acknowledged that ChatGPT metadata was present in four links across two sections of the document. ACCS maintained that the metadata effectively disclosed the use of the tool and said the references had been manually checked. However, efforts to correct the disputed citations reportedly resulted in further inconsistencies. One example involved a research paper that an ACCS source said had been accessed in March 2025. According to the paper’s lead author, however, the research was not published until June 2025, raising additional questions about the accuracy of the report’s references. The report has attracted attention because it was used in the policy process surrounding Australia’s planned restrictions on social media for under-16s. Communications Minister Anika Wells had previously praised the report for identifying potential approaches to age verification. The minister’s department told a Senate inquiry that it had discussed the citation problems with ACCS. However, the department was reportedly informed about faulty links rather than allegations that some references had been fabricated. Australian National University academic Christian Downie warned that unreliable references can have serious consequences regardless of whether they were generated by AI or resulted from human error. He said inaccurate citations could contribute to poor policymaking and weaken public confidence in government-commissioned research. Independent Senator Fatima Payman also compared the controversy with a separate case involving Deloitte, which refunded part of a $440,000 government contract after problems were identified with AI-generated references.

  • |

    Google yanks Nano Banana from Earth after AI images spark chaos

    It took less than a day for Google’s newest AI experiment to become a lesson in how fast image-based misinformation can spread. On Thursday, Google rolled out its Nano Banana 2 model directly inside Google Earth. The feature allowed users to type any prompt and get back a photorealistic image, layered right onto real satellite, aerial, and 3D map data. The intended uses were fairly harmless. People could visualize their dream home or picture how a location might change decades from now. Within just a few hours, users pushed the tool in a very different direction. People began generating fake disasters, staged terrorist attacks, and fabricated sinkholes. Others created images showing refugees or protest crowds in places where nothing of the sort had actually happened. One researcher used the tool to create a fake blast crater in Los Angeles. Another generated an image showing protesters marching right outside Google’s California headquarters. Since the tool used real Google Earth maps as its foundation, these fabricated images looked disturbingly convincing. At first, Google defended the feature. The company pointed to its SynthID watermarking system as a safeguard. It also noted that tools like Gemini or Google Lens could help users check whether an image had been AI-generated. That response did not satisfy critics. Many pointed out that most people encountering a screenshot online rarely take the time to verify it. Critics also argued that the fake images did not even need to stay inside Google Earth to cause harm. Screenshots alone were already spreading fast across social media platforms. By Friday, Google confirmed it was pulling the feature entirely while working on stronger safeguards. The company maintained that every generated image had carried a watermark. It also said none of the images had appeared within the shared, public version of the Earth experience. This incident has reignited larger concerns around generative AI tools and how easily they can be misused, even when protective measures are already built in. It also raises fresh questions about whether watermarking alone is enough to stop confusion once fabricated content starts spreading widely online. For now, Google says its focus is on rebuilding the feature with tighter controls in place. Whether that will be enough to prevent similar situations going forward still remains to be seen.

Leave a Reply

Your email address will not be published. Required fields are marked *