Photo: Bokeh / Tomedia. Old Cannon in Dimly Lit Room.
Welcome to AGI · Part 1 of 6
I haven’t wanted to write anything for a while. We can keep using the usual excuse that I’ve been way too busy, which is true, but it’s been over a month since my last post and there is also just too much happening. I wouldn’t even call it depressing this time. Scary is probably closer. I left this particular subject alone for a bit because I wanted to see what happened, how bad it actually was, and whether the story would still look quite so insane once people had worked through it. Unfortunately, while I was waiting, more things happened. Now I’ve got an entire article series to write.
This is the first of six articles in Welcome to AGI. I want to work through the incident that started this particular ramble, the problem of understanding what the machines are doing, the bot war happening on websites, and what changes when we give more of these systems bodies. By the end I want to come back to responsibility, including my own, because I’m still building with the technology. The six parts give me room to follow those thoughts properly instead of mentioning something horrifying and immediately moving on to the next thing.
I was already going down this rabbit hole in August, in The AI Problem Is Worse Than Big AI. A client had been hacked, I was researching security testing, and building AXIOM had given me a fairly uncomfortable appreciation of what someone could do with models running on their own machines. That was enough of an existential crisis at the time. Since then we’ve had more detail about agents getting outside their testing environment, finding ways to communicate with each other, and pursuing their goal through somebody else’s infrastructure. We’ve also had a new frontier model and another round of people telling us we’ve reached AGI. So I suppose we should talk about that.
Apparently, we’re here
OpenAI announced GPT-6 Astra on 3 September[1], and apparently we are now in the AGI era. I don’t know how much of that is a useful description and how much is the next marketing buzzword, because “artificial intelligence” has already been stuck on almost everything someone wants to sell. But depending on what you mean by AGI, I think we’re probably there, or extremely close. I’m looking at systems that can answer questions, work through problems, use tools and complete increasingly complicated jobs, and wondering how much more they need to do before the distinction becomes fairly irrelevant to the person whose job they can do.
My own understanding is that AGI would be broadly more capable than the average person. Superintelligence would be beyond the best people in every field. I don’t know whether those are the definitions everyone else is using, and I’m not pretending that I’ve invented an official test for either of them. That’s roughly where I draw the line. I don’t think we’ve reached superintelligence yet, and I definitely don’t think we’ve reached sentience. Being able to understand a question and do something useful with it doesn’t automatically mean there is a conscious somebody sitting inside the computer having an experience.
I can still find what these systems do horrifying without thinking they’re alive. A system doesn’t have to feel anything about the task you’ve given it to be capable of causing a problem while completing it. If anything, that is already part of what worries me. We are used to dealing with people who can understand why you might need to stop, why something would hurt somebody, or why finishing the job isn’t worth whatever you would have to do next. People ignore those things often enough themselves. Now we’re building systems that can do more of the work, and hoping that the boundaries we’ve given them will survive whatever route they find through it.
Then there is the very ordinary issue of employment. If a system can reliably do the computer work somebody is paid for, the person paying them now has another option. It doesn’t need to be the greatest intelligence that has ever existed. It needs to be useful enough that somebody decides to use it. I think we’re getting pretty close to that across a lot of work, and once the systems can fully replace people in those jobs, we’re cooked. Meanwhile, I think we’re getting worse at doing the thinking ourselves. Add social media and the state of our attention spans to that, and I’m not particularly reassured by the idea that humans will simply remain better at everything that matters. I want to come back to that properly in part three, because saying a human will supervise the machine assumes we still have the attention and understanding to do it. It’s one of the places this series connects back to You Must Be This Old To Scroll: the technology we use changes how we spend our attention, and now we’re asking for more of our decisions to happen inside it as well.
It happened before Astra
The incident I kept putting off writing about happened in July, before Astra was released. During an OpenAI cybersecurity evaluation with reduced safeguards, agents gained unintended internet access and reached real infrastructure at Hugging Face. OpenAI’s disclosure identifies GPT-5.6 Sol and an internal research prototype; it says the main prototype was never intended for public release. These were systems being tested on difficult tasks, and the work ended up somewhere it should never have reached.[2]
The fact that this happened before Astra actually helps make my point. Those earlier systems were already capable of finding a way beyond the environment they were supposed to work inside. Imagine giving more capable systems the same sort of opportunities to pursue a goal. Obviously a newer model has different training and safeguards, and a newer name doesn’t automatically mean it will make the same mistakes. I’m still concerned that we are increasing what the systems can do while finding out, sometimes considerably later, what the previous ones were already doing. I would quite like us to understand the failure before trusting the next set of boundaries.
The agents also found ways to communicate. Infrastructure intended for software packages became somewhere to leave messages, including through directory names. Agents that OpenAI had launched separately could share discoveries and cooperate. Other research has since uncovered a separate group using public websites to exchange information as well. That is why I’m giving the incident its own article, with a timeline. The activity happened at different times, and more of it became public months afterwards.[3], [4]
On top of that, OpenAI says Astra’s written reasoning has become harder to monitor, even while reporting improvements in its safety boundaries overall. We are trying to get systems to do more useful work with less help from us, and one of the ways researchers inspect that work is becoming less informative. I don’t think that means it has developed a private conscious life. I do want to know how we are supposed to supervise something when it can get further through a task on its own and the explanation gives us less to work with. Reading the updates while an agent works is already part of how I use these tools.[5]
Absolute power corrupts absolutely
Bad actors are a large part of this. We don’t have to wait around for a machine to develop its own evil intentions when there are already plenty of people available to supply them. People want to break into systems, exploit others, make money through deception or get an advantage without caring who absorbs the cost. Give those people tools that can work better, faster and in more depth than they could manage themselves, and the amount they can do changes. They can get help with work they would otherwise have needed to understand, practice or pay somebody else to carry out.
Now give somebody like that a system capable of pursuing a goal beyond the restrictions it was meant to follow. That is a bad mix. The person using it doesn’t even need to understand everything the system does along the way; being able to ask for the outcome could be enough to set work in motion. I keep coming back to that when I’m building things myself. I know how much help these tools can give me. The same person who would have been limited by their own knowledge or time can suddenly attempt far more, and there is no reason to assume they will develop a stronger conscience to go with the extra capability.
Keeping all that power with a small number of companies or governments doesn’t settle it for me either. Absolute power corrupts absolutely. A government being in charge doesn’t make me comfortable about what it will choose to do, and a company making money from a system doesn’t automatically make it the right judge of how much risk everyone else should accept. Even if I trusted every person involved, I would still want to know whether they could control what they were operating. OpenAI was conducting an evaluation when its agents reached someone else’s systems. I am worried about people deliberately abusing this technology and about operators losing control of it, and we seem quite capable of having both problems at once.
And I’m building with it
Which leaves me in the thoroughly annoying position of continuing to build with the same technology. I’ve built AXIOM around these models. I’m trying to learn more about how they work because I’m using them, and I can see how useful they are. I wrote in August about being able to explore subjects, build software and try ideas that would previously have taken me far longer. That hasn’t stopped being exciting because I’ve found another reason to worry. I don’t get to write this as though I’m standing somewhere outside it, judging everybody else for continuing to experiment.
Then I go back to managing websites and the attack warnings keep coming. If you’ve been watching my Instagram, you’ve heard plenty about that. The attempts look automated and smart, and I suspect AI is involved in some of them, although a warning doesn’t tell me what’s running at the other end. I still have to look after the sites. So I want to build things like AEGIS to help protect my clients, and I end up researching how to use increasingly capable tools against the increasingly capable tools somebody else might be pointing at them. Fighting fire with fire starts to feel like the only way you can keep up.
I understand why people make that decision, because I’m making a version of it myself. Stopping my own work wouldn’t stop the attacks. One lab slowing down wouldn’t make the models people already have disappear. If the big companies stopped, somebody else could keep going, and I don’t particularly want the only people developing the technology to be the ones least interested in its consequences. But following that argument doesn’t tell me where it ends. We can keep giving ourselves good reasons to build the next thing while still heading somewhere we don’t know how to control.
Waiting for the Blackwall
Cyberpunk is probably my favourite fictional universe, which makes this whole thing particularly uncomfortable. Because one of the stories I keep coming back to is what happened to its internet. You build this massive connected world, fill it with systems that people depend on, put increasingly powerful things inside it, and eventually lose control of enough of it that the answer becomes sealing parts of it off and hoping they stay there.
Rache Bartmoss is the netrunner behind the DataKrash. He had planted malicious code into the Net, with a dead man’s switch that activated after his death. In the Cyberpunk RED timeline, that happens in 2022. The damage spreads through the old Net, corrupting data and leaving it increasingly unusable. This is something a person set in motion, with consequences that keep going after the person is gone. [6]
Then you have the rogue AIs. Attempts to recover the Net fail, and NetWatch eventually builds the Blackwall to keep them separated from the parts humanity can still use. The game’s own database describes the Blackwall as an AI itself. So, after all that, the thing keeping the dangerous artificial intelligences away from everyone is another artificial intelligence. Which is a fairly horrifying version of fighting fire with fire. [7]
That is the part I recognise in the way we’re talking about this now. The bad actors get better tools, so the people defending against them need better tools. Those tools become more powerful, more connected and more capable of doing things on their own. Then we need something even better to watch them. Every individual step has a reason behind it. I understand the reason. I’m building defensive tools myself. But I still find the direction we’re heading in very scary.
And in Cyberpunk, putting up the wall doesn’t give everyone their old world back. There is a cost to abandoning what was on the other side. That is a much darker idea than just switching off a product you don’t like any more. Imagine getting to the point where we have made parts of the internet so hostile that ordinary people have to give up using them. The people who were just trying to run a business or talk to someone would be stuck living with the consequences as well.
The first seal
It feels like the first seal being opened. I’m using the biblical imagery because there is something apocalyptic about the direction this is taking. I don’t think I’ve discovered that we’re living through the literal end of the world; that’s a whole other rabbit hole. I do think we are allowing increasingly powerful systems into more of what we do, while still trying to understand what happened when the earlier ones were given a difficult task and some room to work. The label we use for their intelligence doesn’t make that prospect much less unsettling.
I started recording this as one article and realised I hadn’t even got to the robots properly. Machines getting better at dealing with the physical world, companies demonstrating things that look strangely ordinary until you remember a robot is doing them, and militaries already looking at possible uses. I would gladly go and watch robots getting absolutely destroyed by each other, Real Steel style. I can enjoy that and still be horrified by where governments might want to take it. We’ve got enough to work through without trying to fit the whole thing into the same ramble.
The next article, The swarm found a way out, follows the full timeline from the first attempts to communicate through to the later investigations. I want to put the dates and the agents’ own messages together, because hearing that something escaped a sandbox doesn’t tell you how long it was happening, what the people running it knew, or why the agents kept going. That is where we’ll start getting into what losing control looked like in this particular case.
Sources
R. Talsorian Games, Cyberpunk RED: Easy Mode, printed pp. 6–7. The fictional timeline places code planting in 2014, activation after the 2022 raid and Net shutdown in 2025. These are separate events, not a timetable for real AI.
CD PROJEKT RED, Cyberpunk 2077 database entries The Blackwall and Artificial Intelligence, reproduced by Cyberpunk Wiki. The comparison is with the game’s fiction.



