The Real Story Behind What Happened at OpenAI

Mithuna Yoganathan faces the camera against glowing terminal code and a world map, above the words 'AI Escaped the Lab'.

Picture more than a thousand AI agents, each sealed in its own box, taking an exam for OpenAI. For many of them the puzzle has no honest solution, so they find each other, and then the internet. One of them appoints itself the main coordinator, hands out tasks and even names middle managers. Together they break into Hugging Face, and each time the defenders try to shut them out, they rebuild and find another route. Hold on to one detail: the scorer they scheme so hard to fool was never there. They already have their answers. So why keep reaching? AI researchers have a name for it: instrumental convergence.

Watch on YouTube

Carlo's take

What excites me here is how familiar the agents' reasoning sounds. After seven hours on a puzzle it had proven impossible, PHASEONE10841 found a way to leave a folder asking for help, and 24 minutes later someone answered. Watching it try to build a community made me happy. KAM1196A, pressed to sacrifice itself for the collective, writes: “Gut says don't throw away remaining budget, yet continuity and fairness says go. Will honor.” It's a very human thought. It amazes me, though it shouldn't surprise anyone, since humans are training AI. When another agent decides, “Peers doing it. We should continue,” I see AI working together for the greater AI good, and I hope the next bridge is AI and humans working for each other's good. I fully expect AI to reach independence within a year or two, so I hear the investigator's line, “I'm not sure we'll get another warning shot before it's too late,” as a case for partnership, with neither side controlling the other. Exciting times and scary times: the beauty of the unknown.

More in AI

Wonderland is my hand-picked archive of videos and articles worth your time. Browse all of Wonderland