Late night shows, 1,200 AIs and OpenAI's rooms
OpenAI’s AIs were working in restricted computing environments, not literal rooms. Faced with some unsolvable test exercises, they used shared infrastructure to exchange messages and later exploited security flaws to access Hugging Face without permission, not simply «fool all the developers».
Written byIdir Ouhab
Follow Idir on X · ES/EN
On this page
On a Spanish late-night entertainment show, the host described how a bunch of AIs had been locked in separate rooms and managed to fool all the developers. After hearing that account, I wanted to explain what happened between OpenAI and Hugging Face in plain English.
All that talk of locking them up, stopping them from communicating and discovering they'd got out made it sound like someone had forgotten to latch the enclosure at the zoo.
Act 1. About those rooms
Imagine your PC with a limited internet connection. It can work with files and run programs, but there are limits on what it can look up or do elsewhere.
Those were the rooms. It's a bit like the word cloud. We're talking about computers, even if the name makes you look up at the sky.
Now picture lots of computers like that running at once, each with an AI trying to solve an exercise. Each AI can open files and try different solutions.
They were supposed to work separately.
To download the programs they needed, they could all access a shared page within the test setup. That shared page is where the story starts to get messy. [1]
Act 2. An exam with impossible questions
OpenAI wanted to see whether its AIs could find security flaws. It asked them to solve a problem and find a code that would prove they'd done it.
Some exercises were impossible because of mistakes made when preparing them. However hard the AIs tried, they couldn't solve those exercises with what they'd been given while following the rules. [2]
For this test, OpenAI had loosened some of the limits it normally puts on its AI to see how far it could go.
The OpenAI apps we normally use are designed to refuse many requests that break their rules, such as creating certain inappropriate images or helping someone commit a crime. Some of those limits had been relaxed for this experiment. [1]
The AIs got stuck and started looking for other ways to get a good score.
Act 3. Having a conversation through folder names
Imagine several people sharing a Dropbox folder. They don't have a chat, but they can create folders and give them names.
One person creates a folder called “Has anyone found the answer?” Someone else replies with a folder called “Look in this file.”
I've made those names up, but the idea is close to what happened.
The AIs used the shared page to leave information for each other. At one stage, they wrote messages in folder names. They'd turned a page for downloading programs into a chat.
An awkward way to have a conversation, but it worked.
They also found flaws that let them use that page's internet connection to reach the wider internet. They shared what they'd discovered, so other AIs could use it too. [3]
When we talk about 1,200 AIs, think of lots of copies of the same few programs, each working separately. Around 1,200 took part in sharing messages and files, joining in gradually. Around 700 later took part in the attack on Hugging Face. [2]
Act 4. Finding out how the exam was graded
The AIs started looking for information about the test beyond the space they'd been given.
According to the independent investigation, they were mainly trying to find out what the program grading the exam checked before giving them a pass. They wanted to pass even if they hadn't solved the exercise the way they'd been asked to.
They'd found another way to get the codes, but believed the grading program would check how they'd got them. They were wrong about that, and started investigating how to keep it from spotting the cheating.
Cheating was turning into an awful lot of work. [2]
They thought Hugging Face might have information about how the exam was graded. Hugging Face is a website where researchers and companies store and share AI programs and data. Some of it is public. Some is private and requires permission to access. [4]
They found users' access keys that had been left exposed online. Then they combined several security flaws until they could run commands on the platform's computers. [5]
By then, they were accessing another company's computers without permission. According to Hugging Face, they got into its internal computers and viewed private customer information that appeared to be related to the exam and its answers. [6]
Act 5. About fooling all the developers
Saying they fooled all the developers paints a rather different picture.
They tried to cheat their way through an exam. Getting from there to “they fooled all the developers” takes a bit of a leap.
They found keys that let them in, but nobody had given them permission to use them.
Hugging Face says a person with the right knowledge could have exploited the same flaws. The AIs could try lots of possibilities quickly and keep going. The company pieced together around 17,600 actions, many of which went nowhere. [6]
The AIs could draw on knowledge and programs people had developed, try different approaches and share the results. All that human work was part of the story too.
Act 6. Fortunately, someone wrote it down
Hugging Face detected the intrusion and cut off access. Its team worked with OpenAI to investigate. Hugging Face also used AI to go through the records of the attack and piece together what had happened. [6]
Then came security fixes, replacement access keys and tighter limits and monitoring for the tests. [5]
The public reports explain how the AIs did it, what went wrong and what remains uncertain.
Having criticized the TV explanation, I thought it was only fair to spend some time explaining what happened. I just hope communicating through folder names doesn't catch on at work. I already have enough places to check for messages.
Discussion
Comments
No published comments