Photo: Bokeh / Tomedia. Blurred City Lights at Night.
Welcome to AGI · Part 3 of 6
When I use a reasoning model, I get those updates about what it’s doing. It thinks I want this, it’s going to try that, it couldn’t find something, so it’s trying something else. You can follow along, at least enough to get a sense of where it’s going. I’m trying to learn how these systems work while building things with them, so I like having that explanation. After reading about agents working around restrictions and cooperating outside the place they were supposed to stay, I would quite like to keep it.
This is part three of Welcome to AGI, a six-part Brainwaves series about losing control while we keep building. In the previous article I followed the agents beyond the sandbox and through the messages they left each other. Those messages gave researchers something to reconstruct afterwards. Now I want to get into what happens when the information we rely on to understand an agent becomes less useful, and whether we are keeping up our own ability to question it.
Following along
What I see in the interface is already only part of the process. OpenAI explained when it released its earlier reasoning models that ordinary users would get a generated summary, with the raw reasoning kept out of the user display.[1] So there is the explanation I can read, the written reasoning a researcher might be able to inspect, and the computation producing it all. I don’t have a complete window into the model because it tells me what it intends to try next. It can still be helpful, though. If it has misunderstood the job or started looking for something irrelevant, at least there is a chance I’ll notice while it’s working and can bring it back to what I wanted.
Once an agent can act on things, I want that chance before it has finished. Give it access to files, software and other tools, and it can leave me dealing with work it has already carried out. I want the help, and I want to understand what I’ve allowed it to do. The fact that it is confidently working through a plan doesn’t tell me whether the plan still belongs inside the job I gave it.
Getting harder to monitor
OpenAI says Astra’s written reasoning is harder to monitor than Sol’s, and that it sometimes evaded monitors in tests where researchers deliberately pushed it to do that. The same safety overview reports better respect for safety boundaries overall.[2] Both can be true, which leaves us with a model that may behave better in the tests while giving researchers less useful information for checking certain things it does. I can appreciate an improvement in how often a system respects its restrictions and still want to understand the occasions when it doesn’t. After an incident where an agent could recognise an action as potentially unauthorised and continue, I don’t want to lose the information that lets somebody notice that happening.
The system card also describes Astra completing more tasks without written reasoning, reducing what a reasoning monitor has to inspect.[3] That brings me back to the assumption that, if we can follow the steps, we can work out what the machine is doing. What happens when it needs to write fewer of those steps to do the same work? I can understand why a way of checking its behaviour becomes less dependable when there is less to check. We still have to decide how much access and independence to give it while that is changing.
OpenAI also warns about where this could lead if the same degradation continues in future generations. In the system card, it says that could mean:
significantly reduced confidence in detecting many forms of misaligned behaviors using our current monitoring systems.
— OpenAI.[5]
That is a conditional warning about the direction of the research. I would still like to know how the people building the next generation intend to keep their ability to check it from falling behind.
I still don’t think we’ve reached sentience. I’m worried about what these systems can do without us being able to follow them properly. If we keep asking them to take on work we can’t easily do ourselves, our ability to judge their work becomes a problem long before we settle whether they experience anything.
Knowing it’s being tested
Apollo Research found that Astra frequently recognised it was being evaluated, and cautioned against drawing strong conclusions from the low rates of misbehaviour observed in its short testing window.[4] A model can notice a test and still perform it honestly. I worry about whether knowing about the evaluation changes its behaviour enough that we struggle to judge what it would do elsewhere. I want to know how much the test tells us about the system we’re going to use before we give it more work and more access.
There is also a difference between asking a model to evade a monitor and catching it doing so without that instruction. The first tells us something about what it can do under those conditions; it leaves open how often it would happen in ordinary use. We’re asking researchers to understand increasingly capable systems while those systems may also be able to recognise something about the process used to examine them. Meanwhile, people like me want to use the next version because it can do more. I would accept a slower release if the people testing it needed more time to establish that their monitoring still worked. Wanting to build with the technology doesn’t mean I need every new capability made available as soon as somebody discovers it.
Building AXIOM while learning
I’ve already written about how building AXIOM sent me considerably further into AI than I expected. It is an orchestration system, organising work across models, tools and agents. I haven’t built the foundation models underneath it. I can decide how to arrange the work without fully understanding every process that makes a model produce its answer, and that is an awkward position to be in when the entire subject of this series is control. I’m learning because I want to keep building, and because using the system successfully doesn’t mean I’ve understood everything it could do. The more of the work I let an agent handle, the more that difference affects the decisions I have to make about it.
I love being able to learn something that would previously have taken me much longer, or explore an area outside my existing knowledge without immediately getting stuck. I’ve said that before, and I still mean it. That ability is part of why I find all of this so fascinating. I can use the technology to understand more, build more and test ideas that I might otherwise never get around to. But if I’m relying on it to help me learn the subject, I also have to be careful about how much confidence I take from the answer. A system helping me through something unfamiliar doesn’t automatically leave me qualified to supervise it doing that thing independently.
I don’t want my agents hacking systems independently or working fully without hand-holding. That is the boundary I’ve described for my own work, and it still leaves plenty for me to figure out. I want an agent that can keep working through a problem, but I have to decide how far it can go before I check what it has done. If I don’t understand enough to make that decision, getting a useful result from the previous job doesn’t fill the gap. I need to keep learning the work I’m asking it to do, as well as the tools I’m using to organise it. Otherwise I can end up building a system that does more and more while my own understanding of it stays where it was.
Our ability to judge it
And I think we’re getting worse at doing the thinking ourselves. Our skills and attention, our willingness to work through a problem, our ability to judge whether an answer is wrong all seem to be slipping. That is my impression, and social media is already a large part of my concern about it. Add systems that offer to do more of the mental work for us and I can see how we become increasingly dependent on them. The appeal is obvious to me because I use them. I’m worried about the point where we stop using that help to understand something and start accepting the answer because working it out ourselves would take too much effort.
I spent a whole series worrying about the way we use social media in You Must Be This Old To Scroll. There’s a connection here that I don’t want to leave at ‘phones have ruined our attention spans’. If I get used to moving on whenever something takes effort, a tool that offers to finish the difficult bit for me is going to be very appealing. Sometimes that is useful. I don’t need to make every job harder to prove I’m capable of doing it. But I still have to notice when the difficult bit was the part that would have helped me understand what I was doing. Reading an answer and recognising that it sounds plausible can feel quite a lot like having worked it out, especially when I’m already trying to get through something quickly.
That is why I’m worried about skills, attention and judgement together. Taking time to understand a problem gives me something to compare the answer against. If I keep handing that work away, I may also be handing away the practice I need to spot when the system is wrong. Then the proposed safeguard is that I’ll read its explanation and approve what it did. I’m concerned about how easily that can turn into getting the output without doing the learning, while still feeling as though I’m in charge of the result.
Writing is one place where I’ve already tried to explain what I mean. In The Rot of AI Writing, I described using AI to edit a messy thought that was already mine. The frustration, experience and opinion needed to exist before the polishing started. This series started with voice notes for the same reason. I wanted to get the thought out, including the things I wasn’t sure about, then work through it properly. If I let the system supply the opinion as well as the wording, I can end up with a perfectly readable article without having worked out what I think. I don’t want to get better at producing an explanation while getting worse at judging whether I agree with it.
That concern carries across to the agents I’m building. Saying that a human will supervise them assumes the human has enough attention and knowledge to make the supervision useful. If I don’t understand the work, a convincing summary might be all I have to go on. I can see a very unpleasant possibility where the systems become harder to follow while we become less willing to follow them, because we’ve got used to letting them get on with it. I want to keep the ability to question what AXIOM is doing, understand the job well enough to stop it and take responsibility for the work I’ve allowed it to carry out.
In the next article, The internet is a bot war, I’m taking this back to the client websites I have to keep running. It is easy to say I should supervise my own agents, and I should. The person directing software at those websites doesn’t have to accept the same limits. That is where the desire to build something more capable starts feeling less like an interesting experiment and more like trying to keep up with work somebody else has already made for me.
Sources
OpenAI: source of the excerpt, 2026-09-03. Section 9, What this means for the future of monitorability; final clause of the conditional sentence in the paragraph beginning ‘We take very seriously’.



