How GPT-6 surprised me

Since I started using GPT-6, delegating work to Codex has begun to feel different.

Previously, I often watched the work in progress and stepped in with more instructions. That is not the right direction. That is not what you should be investigating now. Please reconsider the original assumption. Even after handing over the initial request, I would often need to stay alongside it and adjust its course.

With GPT-6, I feel that this happens less often. When it gets stuck partway through a task, I have seen it step back and reconsider before trying to push ahead. It questions the assumptions it was given and looks for another way forward.

At work, we sometimes discover that our initial assumptions were wrong only after we begin doing the work. Building a screen deepens our understanding of the user, and we realize that the requirements themselves should change. We set out to organize some documents, only to discover that the real problem lies elsewhere. Good work includes this kind of movement back and forth.

Perhaps we can begin to expect the same from AI. What surprised me this time was the change in what I expect from something I entrust with work.

If I can work with a partner that understands the purpose, makes its own judgments, and sometimes even challenges my thinking, I feel I might get closer to the way I have long wanted to work.

The joy of work: sharing principles with capable people and achieving more than we imagined

For me, work is most interesting when someone brings back something better than I had imagined. I, in turn, bring them something that goes beyond their expectations. When we put those contributions together, something emerges that neither of us could have conceived alone. That is the kind of work I want to do again, as often as I can.

This was also what I enjoyed about management. Instead of spelling out every step, I would share where we wanted to go and the principles I wanted us to uphold in our work. I would explain why we were pursuing that goal and what kind of organization we wanted to be, then leave the approach to them. Some people would do wonderful work in ways I would never have thought of.

It is a little like telling someone the destination, the arrival time, and what matters along the way. You do not need to decide which form of transport they should take. They assess the situation, think for themselves, and find their way there. Along the way, they come up with ideas I did not have. I found that inspiring.

I think soccer offers a similar kind of enjoyment. The team shares an understanding of how it wants to play, and each player acts on what they see in the moment. The positions of teammates and the moves of opponents are constantly changing, so the next decision has to be made on the spot. Shared principles are what allow those individual decisions to come together as good play.

I hope for the same kind of working relationship with AI. I want to explain what we are trying to accomplish and what should guide its decisions. Then I want it to choose and combine the work that needs doing, and suggest a better approach if it finds one.

That hope is at the root of my desire to build AI agents that can work autonomously. I want to work together beyond the limits of what I can picture in detail myself.

Not every task needs that level of ability and freedom

At the same time, when I look across our day-to-day operations, some work does not need that much thought every time.

Some tasks have established checks and procedures and can be repeated reliably. To delegate that work, we first need to define it properly so that someone with the necessary skills can carry it out to the required standard. It is natural to consider the right people and costs for the job.

Does someone capable of solving difficult problems need to work out an approach from scratch every time? How much better will the result be because we used that person's abilities? Where the procedure is already settled, the work sometimes goes better when someone simply follows it through reliably.

Of course, even work that usually follows a procedure may require fresh thinking when something unexpected happens or the approach itself needs reconsidering. Within the same job, some parts can follow a procedure and others call for stopping to think. We need to recognize the difference and adjust how we delegate.

Wanting to do the best possible work and applying the highest level of intelligence to every task are two things we need to consider separately.

If I try to apply the way of working I find most rewarding to every task, I can lose sight of what the work actually needs. I think the same thing is happening with AI.

It is time to apply familiar management practices to AI

When we give a highly capable AI the freedom to think, searching for an approach and trying things again both count toward usage. Tokens, the units used to measure the information AI processes, are not consumed only by the final answer.

This becomes a very immediate concern when I think about running agents overnight. I want the work to continue while I sleep. But what if I wake up to find that the usage allowance is exhausted? If the agent goes through more trial and error than expected, how far might the cost rise? I have handed over the work, yet I cannot sleep because I am thinking about how much it is consuming. The more I hope for autonomous work, the more I feel both a reluctance to stop it and anxiety about excessive usage.

In human organizations, we have long adjusted how we delegate to suit the person and the task. When we face a problem whose answer we cannot yet see, we share the background with someone we trust and make time to think together. For recurring work, we establish procedures and decision criteria before handing it over. Even people who can solve difficult problems have limited time, so we think carefully about where to direct their efforts.

We are reaching the point where AI calls for the same judgment.

For problems with no clear answer, provide advanced intelligence and the discretion to reconsider even the underlying assumptions. For work that can be defined as a set of tasks, prepare manuals or workflows and use a lower-cost model that meets the required quality. When a problem falls outside what the procedure can resolve, route it back to a stage that allows for fresh thinking.

Alongside knowing which models are capable, we need to judge how much intelligence and discretion each job requires. I feel that as AI becomes more capable, this kind of management will become a greater part of using it.

The purpose of workflows is changing, too

Previously, getting AI to do work well meant breaking it down into small stages. First gather information, then organize it, then develop a proposal. We took work that did not go well when delegated all at once, divided it into manageable pieces, and put them in order. Workflows were a way to get more out of the model's capabilities.

As models improve, there are more situations in which they can proceed without such detailed instructions. Give them the objective, and they can work out the necessary stages themselves. It is natural to wonder whether we should simply give them more freedom.

But if we keep leaving everything to them, we start to notice the cost of all that thinking. And so we come back around to workflows.

Even when the work leads to the same result, the amount of exploration needed changes depending on whether we ask AI to work out the sequence every time or give it a known procedure. If a procedure has been thoroughly tested, using it can reduce the thinking we repeat.

In this setting, a workflow becomes a way to allocate intelligence and budget to the places where autonomous thinking adds value. By organizing routine stages, we can devote more effort to questioning assumptions and finding new approaches. We should also leave room to return to an earlier stage when a new problem emerges along the way.

Still, splitting work into smaller pieces and passing them to cheaper models will not necessarily reduce the cost. Delegating to another agent means explaining the background. We then need to understand the work it returns, check whether it is correct, and revise it if necessary. Those handoffs and checks take time and tokens, too.

It is possible that, after splitting up the work, we find that one highly capable model could have completed the whole job faster and at a lower cost. What matters is not just the price of the model, but the total cost of reaching a result that meets the required quality. We need to try different ways of delegating each kind of work, check the results, and adjust.

Making an ideal way of working sustainable

Capable people share a purpose and principles, act on their own judgment, and bring back work that exceeds everyone's expectations. I have long wanted to create an environment where we can work that way.

I have experienced how rewarding it can be. But I cannot say I have managed to create that environment consistently and sustain it. Finding people with the abilities we need, working together, learning to understand each other, and providing work that lets those abilities flourish: there have been many difficulties in turning the ideal into reality.

Watching AI evolve, I see the possibility of moving closer to that ideal. When I get stuck, it might offer a different perspective. As it takes the work forward, it might return with discoveries I had not expected. I have great hopes for working with a partner like that.

To keep putting those abilities to use, we also need to design how we delegate. We need to organize the parts that can run as routine work and identify where judgment is needed. Questions with no clear answers and work that creates new value deserve enough room to think, and enough budget to support it.

Perhaps it is enough to start with just me and an AI as my working partner. We receive each other's work, and it sparks the next idea. Little by little, we reach places I could not have reached alone.

That is the way of working I find myself thinking about after being surprised by GPT-6. I want work that exceeds our imagination to become something we can do every day, rather than an experience we have only once.


This article was reconstructed from a conversation between @theaktky and Kijitra's Ikeda.