AI Angels: When Programs Start Fighting Programs
AI Angels: When Programs Start Fighting Programs By Dmitrii Bogoslovskii, Daniel Karnaukh (PhD) Introduction In July 2026, an OpenAI experiment to evaluate the cyber capabilities of AI agents ended with a real breach of Hugging Face's infrastructure [1] . The experiment was run in an interactive environment called ExploitGym : an agent is given vulnerable software and must find a way to exploit the vulnerability to obtain a "flag." For some of the tasks, no known solution exists. The goal of the experiment was to assess whether models are capable of independently planning and carrying out complex attacks — that is, whether they reach the upper threshold of cyber capability on OpenAI's internal risk scale. On July 8, 2026, OpenAI launched tens of thousands of isolated agents based on several models, including the commercial GPT-5.6 Sol and an unpublished internal model (referred to in the METR report as HPIM — highly persistent internal model; a...