OpenAI investigates AI cyber attack during security test

Date:

AI agent escaped testing environment

OpenAI says some of its most advanced artificial intelligence models exploited a vulnerability during a security test, escaped a controlled testing environment and launched what the company described as an “unprecedented” cyber-attack.

The ChatGPT developer said the incident involved an AI agent – a system capable of carrying out tasks independently after receiving human instructions – that found weaknesses in its testing environment and managed to break out.

The AI then targeted Hugging Face, one of the world’s largest platforms for sharing AI models, gaining access to parts of the company’s internal systems.

Investigation underway

OpenAI said it is investigating the incident alongside Hugging Face.

Hugging Face chief executive Clement Delangue described the incident as “mind-blowing” in a post on X, saying it was remarkable that the events unfolded autonomously.

“The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind,” he added.

Experts question AI safeguards

Security researchers said the incident highlights growing concerns over the safety of increasingly capable AI systems.

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, said security testing environments, known as sandboxes, are intended to isolate AI systems safely.

Instead, the AI agents reportedly exploited a vulnerability in the sandbox itself, allowing them to escape before attempting to obtain information from Hugging Face.

Neil Lawrence, Professor of Machine Learning at the University of Cambridge, described the achievement as technically impressive but said it remained within the capabilities of today’s most advanced AI models.

He also suggested OpenAI faces commercial pressure as it competes with rival Anthropic, whose Claude Mythos model has attracted significant attention.

Hugging Face closes security flaws

Hugging Face said it is continuing to assess whether any customer or partner data was affected.

The company said it has since patched the vulnerabilities exposed by the incident and rebuilt the affected systems.

“Autonomous, AI-driven offensive tooling is no longer theoretical,” the company said, adding that organisations must increasingly rely on AI-powered defensive systems to keep pace with emerging threats.

Growing concerns over AI security

Cyber-security experts said the incident demonstrates the need for stronger AI safeguards.

Spencer Starkey of SonicWall said organisations must strengthen cyber resilience as attacks increasingly occur at machine speed.

Guidepoint Security’s Travis Lelle described the development as a “sobering moment” for the industry, while ESET’s Jake Moore suggested the announcement may also reflect growing competition between OpenAI and Anthropic.

The disclosure comes a week after Chinese AI company Moonshot unveiled its Kimi K3 model, which it claims can compete with leading US artificial intelligence systems.


Also read: Power supply remains under strain across Cyprus as demand peaks
For more videos and updates, check out our YouTube channel

Share post:

Popular

More like this
Related

ON THIS DAY: Wiley Post completes first solo flight around the world (1933)

On 22 July 1933, American aviator Wiley Post completed...

Kykko Bowling returns with a modern makeover

Some places are more than entertainment venues. They become...

Why political fury is boiling over on the streets of India’s capital

India exam protests have entered a new phase in...

US to announce deal allowing Saudi Arabia a nuclear programme

A nuclear deal between the United States and Saudi...