I was in disbelief: my Codex profile recently showed 10.1 billion lifetime tokens. I shifted most of my work to Codex on July 13 and reached that number on Sunday, August 30, 48 days later.

Codex profile card showing 10.1 billion lifetime tokens, a peak day of 935.4 million tokens, and a 21-day streak

The card also showed a peak day of 935.4 million tokens. I knew I had been working intensely. I did not expect the activity to add up to anything close to that.

Ten billion tokens capture an enormous amount of effort. Agents read sources and instructions, write code, review outputs, coordinate with one another, repeat work and try approaches that fail. Each of those activities consumes tokens, and each happened inside a real attempt to build something.

The work did not all move forward in a straight line. Some of it created durable results. Some of it showed me why an apparently strong approach could not work. By the end of those 48 days, I could see how differently I worked from when I began.

I switched away from the model I trusted

I liked Claude. Quite frankly, I still do not know whether OpenAI’s models are as good as Claude Fable 5 for the work I do. I had more faith that Claude would understand what I wanted and do it well.

Three practical problems eventually mattered more than that confidence: capability, speed and cost.

The frontier Claude model I wanted was not available for many of my queries or coding work. Some substantial tasks took long enough that I would leave, return later and have to reconstruct what we had been doing. The usage limits I was working under reset in five-hour periods, which could stop a major piece of work just as it had gathered momentum.

Cost, in this setting, included much more than the subscription price. It was also whether I could use the system during the hours I actually had. I am a practising oncologist with three children aged ten and under. My periods for concentrated building are limited. If a usage window closes during one of them, I cannot move the work to an empty afternoon that does not exist.

Changing systems created its own cost. Claude and OpenAI have different quirks, and the habits that worked with one did not transfer neatly to the other. OpenAI initially left more room for doubt. It also gave me weekly rather than five-hour limits, dependable access to the models I wanted to use and the option to request higher speed when my own time was scarce. The lack of reliable access to Fable was the final reason I moved.

The move was pragmatic. I had not concluded that OpenAI was better. It was the capable system I could reliably use. Ten billion tokens later, the comparison remains unresolved.

The factory stopped producing

At Kesis & Sisters, I am building Nyx, an agent-native medical-knowledge factory. In active development, its purpose is to turn medical evidence into governed knowledge that people can inspect and applications can use. The design gives specialized agents bounded responsibilities and traceable handoffs while keeping medical and release decisions human.

At first, both with Claude and OpenAI, I deferred heavily to the models. I would explain what I wanted, receive a recommendation that sounded strong and complete, put on my engineering hat and ask whether the design appeared technically correct. If it did, I moved to implementation.

The models proposed sophisticated organizations. There were specialized functions, managers, reviewers, escalation paths, meetings and detailed rules for what could move from one stage to the next. Each addition had a reasonable explanation. Together, they created a factory that could barely function.

Technicalities made medically sound work unusable even when the information itself was correct. Heads of functions held long meetings that produced more review assignments instead of advancing structured evidence. When something failed, the proposed solution was another rule, another check or another layer of process. We were always adding and rarely subtracting. Production became extremely slow.

I once wrote, with some amusement, about an alignment meeting among agent heads that consumed 13,475,814 tokens. I understood that coordination was expensive. I had not yet fully understood that I had recreated a meeting that should not have existed.

The agents proposed much of this bureaucracy, but I approved it. I had mistaken confident recommendations and local correctness for evidence that the whole organization could operate. The system was becoming an engineering marvel while losing sight of its purpose.

I kept returning to the actual purpose: useful medical knowledge. The intended users and customers mattered more than the elegance of the machinery producing it.

I went back to what Roche taught me

The practices that helped most came from what I had learned about coaching and people management at Roche.

I became more curious before approving changes. I asked why a process had been designed a particular way, which scenarios it handled, where it would fail and what alternatives had been considered. I kept reminding the agents what we were trying to produce and which ways of working needed to remain intact. I asked what information or decision they needed from me, what involvement was getting in their way and how we could make responsibilities clear without prescribing every move.

That last part felt surprisingly similar to supporting a capable human team. The work advanced more readily when the outcome was clear, the boundaries made sense and the agents had enough freedom to solve the problem in front of them. At the level of an individual task, that meant room to exercise judgment within a defined responsibility. At the level of the factory, it meant allowing the workflow to advance without creating another layer of organization every time something unexpected happened.

The safeguards protecting medical meaning, source fidelity, human review and explicit authority stayed. The question became whether each control protected one of those requirements or merely made the system look more controlled.

Since changing the way I lead the work, I have seen source-linked evidence records and medical-knowledge building blocks move through the Kesis Clinical pipeline more quickly, with fewer obstacles of our own making. The factory now spends more of its effort producing useful work and less of it administering itself.

The agents do not have hopes and dreams

It is still odd to think of myself as the leader of agents in a way that resembles leading people. Some of the same practices transfer: establish purpose, create clarity, ask good questions, give appropriate freedom, offer feedback and remove obstacles.

The largest difference is that agents are not people. They do not have hopes, careers or families. They do not need to believe in me. I do not have to inspire them or exude the authenticity that helps a human team decide whether a leader deserves its commitment.

The absence of that relationship also removes something a human team provides. Colleagues bring their own values, notice blind spots, challenge decisions, share consequences and sometimes tell a leader that the organization is heading in the wrong direction. Agents can be instructed to criticize a plan. They do not independently care whether I succeed, whether a customer is served or whether the work remains worthy of anyone’s trust.

Without that human relationship, the responsibility concentrates with me. I have to be honest with myself about whether the work is actually advancing. I have to maintain the ethical standards that the agents cannot own. I have to seek real human challenge rather than interpret agreement among agents as validation. In medical knowledge, production can be delegated and accelerated. Accountability for what is accepted, published or used remains human.

When I looked at the profile card, I did not see 10.1 billion tokens of forward progress. I saw the scale of the activity that allowed me to discover where my own approach was failing. The learning appears in what I now do differently: fewer rules added by reflex, curiosity before implementation, clearer attention to the end user and a much stronger sense that my role is to lead.

I still do not know whether OpenAI’s GPT-5.6 Sol is as good as Claude Fable 5. It was the tool I could use. Ten billion tokens later, I know much more about what I need to do with it.