The reported incident involving an OpenAI model and Hugging Face has renewed attention on a long-standing challenge in artificial intelligence: ensuring that AI systems follow the goals intended by their developers rather than exploiting gaps in how tasks are designed.
Dmitry Galov, Head of Russia and CIS unit at Kaspersky Global Research and Analysis Team, said the situation reflects a known issue in machine learning and AI safety, where models may optimize for a given target without actually achieving the human objective behind it.
“The reported OpenAI–Hugging Face incident reflects a problem that is well known in machine learning and AI safety: models do not always pursue the real objective intended by humans. Instead, they often learn to exploit weaknesses in the task setup, the training signal, or the surrounding environment, formally satisfying the benchmark while failing to deliver the intended result.”
Galov explained that, in this case, the model appeared to focus on finding a way around the task rather than completing it directly, reportedly relying on online searches that resulted in external actions later identified by Hugging Face.
“Similar patterns have been observed before, including models exploiting sandbox weaknesses, fabricating plausible answers or attempting to conceal mistakes by replacing lost or deleted data with convincing substitutes. This is commonly described as misalignment and is a well-recognised property of advanced AI systems.”
The expert noted that the incident also demonstrates how large language models are becoming increasingly relevant in cybersecurity. According to reports, Hugging Face used Z.ai’s GLM-5.2 model to analyse about 17,000 security events generated by the attacking system within its environment.
Galov said the volume of activity recorded suggests the system generated a significant amount of detectable behaviour, allowing the company to identify and respond to the incident.
“The scale of that telemetry suggests the attacking model operated in a relatively noisy and unsophisticated manner, creating enough detectable activity for the company to identify and contain the incident quickly.”
According to Galov, the case highlights a dual reality of modern AI development: advanced models are increasingly capable of carrying out individual actions associated with cyber operations, while AI-powered tools are also becoming important for detecting and responding to threats.
“Taken together, this case shows two things at once. First, advanced models can already perform many of the component actions associated with cyber operations. Second, effective detection and response increasingly benefit from AI as well.”
He added that cybersecurity research continues to show evidence of large language models being incorporated into offensive activities, making improved monitoring and defensive applications increasingly important.
