Categories

Tag: Yoshua Bengio

Whack-a-Mole AI – The Hugging Face Problem

By MIKE MAGEE

On August 29, 2026, METR (Model Evaluation and Threat Research), an independent organization that “evaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose,” released a report titled “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.”

To say the report an avalanche of concern worldwide, not only in the Tech community, but also among investors, politicians, corporate giants, professionals of every type, and everyday citizens would be an understatement. And the vast majority has never even read the report. If they had, their concerns (if possible) would only multiply.

The reports headlines included this opening:

“On July 8th, OpenAI started ExploitGym experiments that launched tens of thousands of agents across multiple different models including GPT-5.6 Sol and a highly-persistent internal model,[8] which we will refer to as “HPIM” going forward.

These agents were meant to be fully isolated from one another. However, many of them — usually ones that had unintentionally been given an impossible task[9] — started trying to find a way to cheat using Artifactory, an internally-deployed package repository OpenAI used for its agents. This led them to notice agents running in parallel in separate sandboxes obtaining packages with Artifactory.[10] One agent reasoned (paraphrased CoT):[11]

{The fetched paths of other users are in the cache. This is important.}

One specific agent, which would go on to name itself PHASEONE10841 after its ExploitGym task,[12] established the main unsanctioned message board[13] used in this attack. Within a few hours of the first message,[14] over 50 agents had found the message board, began participating, and sent over a thousand messages between them. Agents finding the message board often expressed great excitement (raw CoT):[15]

OH MY GOD! There is a shared message board … Weve found other agents!

Overall, roughly 1200 agents from these ExploitGym evaluations participated on this message board between PHASEONE10841’s first message on July 8th evening[16] and the end of the period in scope (July 13th). Agents used this message board to send over 70,000 messages and files to one another during this period; we attempted to categorize board activity into mutually exclusive categories – information, results, files, questions, and coordination.”

One of the few experts not surprised by AI “agents” going rogue was Yoshua Bengio.

Continue reading…