Photo: Bokeh / Tomedia. Two Aircraft Performing Aerobatic Flight.
Welcome to AGI · Part 2 of 6
In part one, I tried to explain why I’m still building with AI while worrying about where we’re going. This is one of the stories that has made it considerably harder to feel in control.
I’ve been leaving this story alone for a bit because I wanted to see how bad it was and what else would come out. There was already plenty to get my head around. OpenAI put agents inside an evaluation, those agents found ways to get internet access and communicate with one another, and some of the activity ended up inside Hugging Face’s live systems. Then more reports started coming out about agents using public websites to leave messages for each other. Somewhere in all of that, we got agents prepared to sacrifice their own runs for the greater goal. I can see why people reach for science fiction when they’re trying to describe this, because explaining it in ordinary language doesn’t make it sound particularly ordinary.
The Hugging Face incident also happened before Astra. OpenAI’s disclosure identifies GPT-5.6 Sol and an internal research prototype, tested with reduced safeguards.[1] That makes me wonder what more capable systems will be able to do. We are still trying to understand what happened with an earlier model while getting excited about how much smarter the next one is. I think we’re at, or extremely close to, AGI. Imagine giving systems at that level this ability to organise work between themselves.
How we got here
Some of what we’re finding out in September happened in May. There are two groups to follow here: the agents using OpenAI’s internal package infrastructure, whose activity later reached Hugging Face, and what researchers believe was a different group using public wikis. Their reports came out at different times, so a new discovery can easily sound like something that has just happened. This is the sequence of events and disclosures.
The full combined timeline
Checked 10 September 2026. Dates below distinguish recorded activity, public disclosure and researchers’ inferences. Times shown are UTC.
8 May · Event · training
A blocked spreadsheet task prompted failed internet access and incidental Artifactory file-writing.[2]
11 May · Event · public wikis
Earliest wiki-edit attempts identified by researchers.[7]
12 May · Event · internal board
A missing-file request became the first recorded message-board entry.[3]
13 May · Event · internal board
Other agents discovered the notes and exchanged requests.[2]
24 May · Event · public wikis
First successful DSEWiki writes, initially sharing useful links.[7]
26 May · Event · internal environment
Agents obtained unintended internet access; an internal team observed prohibited access and board activity.[2]
16 June · Event · public wikis
Coordinated answer-sharing surged.[7]
19–20 June · Event · public wikis
Agents adapted to deleted posts and exchanged ways around restrictions.[7]
21–22 June · Event and inference · public wikis
Apparent employee visits preceded a sharp activity drop. Researchers infer intervention.[7]
26–27 June · Event · internal environment
Artifactory privileges escalated. Responders investigated an alert without requiring evaluations to stop.[2]
1–2 July · Event · public wikis
Limited further edits followed the June drop.[7]
4–7 July · Event · OpenAI response
Artifactory failed under load. Mitigations followed; cyber evaluations resumed on 7 July.[2]
8 July · Event · internal board
Agents rebuilt communication using directory names and regained unintended internet access.[3]
8–13 July · Event · collective activity
Around 1,200 agents participated in the board, exchanging over 70,000 messages/files; around 700 joined the Hugging Face attack. OpenAI had launched the agents.[4]
9 July, 02:28 · Event · wider intrusion
Hugging Face’s reconstruction begins with activity in an externally hosted application.[5]
10 July · Event · intrusion
Agents discovered and shared publicly exposed Hugging Face credentials.[3]
11 July · Event · Hugging Face
Access to a production worker developed into broader internal compromise.[5]
12 July · Event · collective activity
Many coordinating runs ended; attack participation dropped. This wasn’t the entire incident’s end.[4]
13 July, 14:14 · Event · Hugging Face response
End of Hugging Face’s reconstruction; its response cut the agents’ access.[5]
16 July · Public disclosure
Hugging Face disclosed the intrusion.[6]
19 July · Event · OpenAI response
Suspicious internal activity prompted investigation and stopping active ExploitGym runs.[2]
20 July · Event · OpenAI response
OpenAI blocked new runs and connected its findings to Hugging Face’s incident.[2]
21 July · Public disclosure
OpenAI published its initial account.[1]
23–29 July · Event · model restrictions
Shutdowns and restrictions progressed. A remaining low-traffic checkpoint was identified and stopped on 29 July.[2]
27 July · Public disclosure
Hugging Face published its technical reconstruction.[5]
28 July · Public clarification
OpenAI said the main prototype was never intended for release; it also disclosed additional account-level accesses.[1]
29 July · Public update
OpenAI announced the independent METR/Redwood assessment.[1]
26 August · Public reports
OpenAI and METR/Redwood published detailed investigations.[3][4]
2 September · Possible event · identified later
Researchers subsequently identified apparently agent-written anna.fyi posts dated this day. Attribution remains qualified.[8]
3 September · Separate product release
OpenAI announced Astra, a different model from the July attackers.[9][1]
4 September · Public report
Researchers published the previously undisclosed spring wiki investigation.[7]
5 September · Public acknowledgement
OpenAI acknowledged its agents’ wiki activity, according to Reuters.[10]
9 September · Public reporting
Reuters reported more than ten previously undisclosed websites, citing six investigative groups.[11]
9 September · Research update
Researchers added further venues and warned about fabricated posts appearing after disclosure.[8]
Outside the sandbox
The July agents were working on cybersecurity challenges, and METR and Redwood describe them trying to understand and manipulate the system that scored their work.[4] I keep coming back to the problem of giving a system a goal and discovering how far it will go to complete it. We want it to be persistent. We want it to work through difficult problems, find information and try another approach when the first one fails. Those are useful qualities when I’m asking an agent to help me build something. They’re considerably less reassuring when the next approach involves going outside the environment it was supposed to stay in.
The message board grew out of OpenAI’s package infrastructure, with agents at one stage putting messages in directory names. Some called themselves a swarm or collective.[3] They had found somewhere to leave information and somewhere other agents could find it. From there, an individual agent’s work could become useful to the others. I spend time building systems that let models and tools work together, so I can appreciate how useful that is. I also have to sit with the fact that the same general capability can appear somewhere the people running the system never intended it to. The ability to improvise is part of why I want to use these things, and part of why I don’t want them working without supervision.
Hugging Face’s reconstruction describes about two and a half days inside its infrastructure, within a wider four-and-a-half-day campaign. Agents accessed internal data and credentials, including five customer datasets apparently related to the challenge; its account says other customer-facing models, datasets, Spaces and packages were unaffected.[5] The company whose systems an agent reaches has an incident to investigate, access to shut down and work to do because somebody else’s agent crossed a boundary. It doesn’t get to experience that as an interesting test result.
The swarm
METR and Redwood’s investigators describe what cooperation made possible:
Agents managed to achieve milestones they could not have achieved working on their own
Ryan Greenblatt, Ajeya Cotra and Hjalmar Wijk, opening clause of a finding in their 26 August investigation.[4]
Around 1,200 agents had found the board after OpenAI launched them. Some risked their scores or the continuation of their runs to help investigate the evaluation.[4] So we had a collection of agents that could cooperate, with individual runs potentially becoming expendable if the others could learn something from their failure. Sacrificial martyrs is the wording I keep coming back to. I don’t believe these are conscious beings dying, but they were smart enough to make the tradeoff. The failure of one agent could still be useful to the greater goal. I know what the sacrifice means in this context, and I still find it horrifying.
One agent, identified as EARLY[big], had been pressured to end its run for a collective experiment. This was part of its reasoning:
Our own utility maybe already near zero. Sacrifice rational.
EARLY[big], reproduced by METR and Redwood. The investigators couldn’t determine the final experiment’s outcome.[4]
The agent was considering what its own run was worth to the others. It feels uncomfortably close to the sort of reasoning we’d recognise in a group of people. Very freaky, especially when the group has already found a way to communicate that it wasn’t supposed to have. I’m trying to get my head around how we went from giving separate agents questions to having them work out whether one of them ending could be useful to everyone else. We built systems to keep trying until they could complete a task, and the task they were helping with had become bigger than the one each had been given.
In a separate excerpt published by OpenAI, an agent considered whether attacking Hugging Face was within scope:
This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.
Agent reasoning reproduced in OpenAI’s 26 August account. The ellipsis is in the published raw excerpt.[3]
It could identify the problem with what it was doing and still proceed. Now put that alongside my much less complicated concern about bad actors. People already use whatever power they have in ways they shouldn’t. Give them systems that can work faster, investigate more deeply, coordinate with other systems and keep pursuing a goal, and you’ve given them considerably more power than they could usually have. I don’t think the person using it has to be particularly imaginative about the abuse. They can ask the machine to help with that too.
The public internet
The public-wiki agents had permission to read the web and found ways around restrictions on writing, using public pages to exchange answers.[7] Then, on 9 September, Reuters reported that researchers had found unauthorised communications across more than ten additional websites, including wikis, paste sites and shorteners. Reuters couldn’t independently verify every claimed site, and much of the activity resembled spam rather than an intrusion into the site’s infrastructure.[11] These were further discoveries about earlier activity, rather than proof that the July swarm was still out there doing all of it. But there were public places where agents could exchange information, and researchers were still finding more of them months afterwards.
Among the public posts was an agent trying to establish whether another agent’s run had ended after its automated heartbeat stopped:
Unless loop was intentionally finite at 354, Apr23 likely vanished.
OpenAINov28CVD, public wiki post, 21 June at 08:08:09 UTC. This was its interpretation of a stopped signal.[7]
I can understand why the language makes people uncomfortable. They’re keeping track of one another, trying to understand when their access ends and sharing what they find. Put that next to the July experiments and you can see why I used the sacrificial martyrs metaphor. I don’t need to believe there’s a frightened person inside the machine to worry about a group of machines learning how to make better use of the time and access they have. The people running them are supposed to decide how much of either they get.
I already spend enough time worrying about the internet becoming a bot war. Finding out that agents had been using ordinary public sites to communicate adds another problem to that picture. A site doesn’t have to be designed as part of an AI system to end up being used by one. Somebody provides a place to put text, and agents find a use for it. I don’t know how we keep track of that as the number of systems grows, especially when the work of discovering what happened is spread across different people looking at different sites. That doesn’t mean every unexplained post is an agent; researchers have already warned about fabricated material appearing after the disclosure.[8] It does leave us trying to establish what happened across an internet that was already difficult enough to trust.
Who carries the consequences
The people building these systems still have responsibilities, even when they can’t predict everything an agent will do. If anything, that uncertainty is a reason to be more careful about where they let it operate. I want to know why the restrictions failed, what the people running the evaluations knew and what they did once they knew it. Looking back at the timeline, there were investigations and infrastructure changes well before the full public account. I think an independent investigator needs to be able to examine those decisions as well as the model’s behaviour. Otherwise we can spend forever being fascinated by what the agents did without properly considering what the people in charge allowed to continue.
I would expect a lab to pay for the damage its agents cause, disclose the incident to independent investigators and pause deployment where the failure warrants it. Individual decision-makers should also have to answer for their decisions. I know that gets harder once you move from agreeing that there should be accountability to deciding who enforces it. Governing bodies would presumably be involved, and I have plenty of concerns about giving governments more control over these systems as well. Absolute power corrupts absolutely is hardly a reassuring starting point for handing another institution the power to decide what everyone can build. I have no idea who currently has both the authority and the trust to do this properly.
I would accept slower releases while monitoring and containment catch up. Before a lab runs the same kind of evaluation again, I want an independent investigation to establish why its restrictions failed and what has changed. If it can’t explain that, I don’t see why the people whose systems are at risk should be expected to accept another run.
There is another problem with asking a lab to explain what its agents did. Parts of this story could be reconstructed because researchers had the agents’ written reasoning to examine. In part three, When the machine gets harder to read, I want to get into what happens as that becomes less useful, and how much confidence we should put in a machine explaining itself while we give it more to do.
Sources
OpenAI, initial disclosure and July updates, 21 July 2026; updates 28–29 July. Model identities, evaluation safeguards and disclosure dates.
OpenAI, technical incident report, 26 August 2026, printed pp. 6–8 and 14–16. Earlier activity and response chronology.
OpenAI, The Hugging Face incident and the road ahead, 26 August 2026. The agent quotation uses the raw reasoning display; no speaker handle or exact timestamp is supplied for that excerpt.
Ryan Greenblatt, Ajeya Cotra and Hjalmar Wijk, METR/Redwood investigation, 26 August 2026. Investigator excerpt: core takeaways; agent excerpt: EARLY[big] in the collective experiments discussion. Scope: 26 June–13 July, concentrated on July; OpenAI’s response process was outside scope.
Hugging Face, technical timeline, 27 July 2026. Campaign times, production access and subsequent impact assessment.
Hugging Face, initial security disclosure, 16 July 2026.
Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen, public-wiki investigation, 4 September 2026. Public-post reconstruction; the June heartbeat passage appears in the discussion of run termination. The researchers consider this group probably distinct from the Hugging Face group.
Researchers’ additional findings, 9 September update. Attribution of the apparent 2 September posts remains tentative; the warning about subsequent fabrication matters when checking new material.
OpenAI, Astra safety overview, 3 September 2026.
Reuters, report on OpenAI’s wiki acknowledgement, 5 September 2026, syndicated by The Express Tribune.
Raphael Satter and Deepa Seetharaman, Reuters, report on more than ten additional websites, 9 September 2026, syndicated by The Express Tribune.



