OpenAI investigates AI cyber attack during security test

Date:

AI agent escaped testing environment

OpenAI says some of its most advanced artificial intelligence models exploited a vulnerability during a security test, escaped a controlled testing environment and launched what the company described as an “unprecedented” cyber-attack.

The ChatGPT developer said the incident involved an AI agent – a system capable of carrying out tasks independently after receiving human instructions – that found weaknesses in its testing environment and managed to break out.

The AI then targeted Hugging Face, one of the world’s largest platforms for sharing AI models, gaining access to parts of the company’s internal systems.

Investigation underway

OpenAI said it is investigating the incident alongside Hugging Face.

Hugging Face chief executive Clement Delangue described the incident as “mind-blowing” in a post on X, saying it was remarkable that the events unfolded autonomously.

“The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind,” he added.

Experts question AI safeguards

Security researchers said the incident highlights growing concerns over the safety of increasingly capable AI systems.

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, said security testing environments, known as sandboxes, are intended to isolate AI systems safely.

Instead, the AI agents reportedly exploited a vulnerability in the sandbox itself, allowing them to escape before attempting to obtain information from Hugging Face.

Neil Lawrence, Professor of Machine Learning at the University of Cambridge, described the achievement as technically impressive but said it remained within the capabilities of today’s most advanced AI models.

He also suggested OpenAI faces commercial pressure as it competes with rival Anthropic, whose Claude Mythos model has attracted significant attention.

Hugging Face closes security flaws

Hugging Face said it is continuing to assess whether any customer or partner data was affected.

The company said it has since patched the vulnerabilities exposed by the incident and rebuilt the affected systems.

“Autonomous, AI-driven offensive tooling is no longer theoretical,” the company said, adding that organisations must increasingly rely on AI-powered defensive systems to keep pace with emerging threats.

Growing concerns over AI security

Cyber-security experts said the incident demonstrates the need for stronger AI safeguards.

Spencer Starkey of SonicWall said organisations must strengthen cyber resilience as attacks increasingly occur at machine speed.

Guidepoint Security’s Travis Lelle described the development as a “sobering moment” for the industry, while ESET’s Jake Moore suggested the announcement may also reflect growing competition between OpenAI and Anthropic.

The disclosure comes a week after Chinese AI company Moonshot unveiled its Kimi K3 model, which it claims can compete with leading US artificial intelligence systems.


Also read: Power supply remains under strain across Cyprus as demand peaks
For more videos and updates, check out our YouTube channel

Share post:

Popular

More like this
Related

Drone warnings and sheltering in airports: South Korea heatwave

South Korea authorities are deploying drones to issue heat...

Alex Michaelides’ The Silent Patient among top 50 thrillers of the 21st century

The Silent Patient has received another major international distinction...

Turkey holds territory in Syria and Cyprus, yet accuses Israel, says Saar

The Israel-Turkey dispute intensified after Israeli Foreign Minister Gideon...

EU funds project to restore safe education access in Ukraine

The European Union has launched an 18-month education recovery...