I seem to have built a small company out of AI agents.
It has specialists, managers, departments, independent reviewers and enough process to make everyone feel properly supervised. It has no office, payroll or holiday party. So far, morale appears stable.
Two things happened inside this company that have been difficult to unsee. The management layer needed less reasoning than the workers. The department-head meeting recorded 13,475,814 tokens.
The manager was set to Medium
I first noticed the hierarchy while configuring a workflow for data extraction.
The agents doing the extraction had the difficult job. They had to read the source, understand it and produce structured information without dropping the details that mattered. I gave them the XHigh reasoning setting.
The orchestrators ran best at Medium.
Their job was to assign the work, keep track of what came back, return incomplete results and decide when the next step could begin. They needed to understand the contract for the work. They did not need to perform the extraction themselves.
Medium and XHigh describe how much reasoning the model uses. No annual performance review was involved. Still, I resisted drawing conclusions about management for almost three seconds.
Before anyone updates LinkedIn, a good human manager does much more than these orchestrators. A person hires, persuades, resolves conflict, notices when someone is struggling and accepts responsibility when the plan fails. The orchestrators had no one to motivate and no budget to defend. They kept the queue moving and remembered the rules. They were very good at both.
The arrangement made an uncomfortable amount of sense. The specialists needed the largest reasoning budget because they were doing the hardest cognitive work. Their orchestrators needed clear instructions, dependable memory and the willingness to send something back.
The biggest reasoning budget did not have to go to the boss.
The alignment meeting ran for 13 million tokens
Then my artificial company discovered cross-functional alignment.
This workflow brought together Medical Affairs, Technical, Quality Assurance and Data Science. The four functions worked as peers, each testing shared work against its own standards. Technical, for example, checked whether work from Medical Affairs met the technical requirements. A finding sent the work back. A correction returned for another review. Everyone had to be looking at the same version.
This was the AI equivalent of four department heads sitting around one table. Each arrived with a different reason to object.
That table stayed occupied for 13 hours and 13,475,814 tokens.
The total excluded the actual extraction runs and every minute spent waiting for a human. It covered the conversation around the work: reading another function’s output, challenging it, responding to findings, checking corrections, confirming the current version and deciding whether the next review could begin.
I had automated the alignment meeting and discovered that the alignment meeting was enormous.
In a human organization, much of this would happen through a meeting, a pre-read, several reply-all emails and possibly a smaller meeting before the next meeting. My artificial department heads needed no calendar. Nobody was double-booked. Nobody asked to move the discussion to Thursday. At no point did anyone say, “Can everyone see my screen?”
Thirteen hours is still a very long meeting. It did not require four human department heads to find the same open space on four calendars.
I do not have a corresponding API bill. The workflow did not run through the API, so any dollar figure would be inferred. The token count still shows how large the coordination layer was. Checking one another’s work was a substantial body of work on its own.
Watching the agents made alignment less mysterious. It was a pile of small, exact questions. Is this complete? Does the representation preserve the medical meaning? Does it meet the technical standard? Did the reviewer examine the current version? Has anyone exceeded their authority?
Humans often answer those questions in meetings because meetings are where the relevant people can be assembled. The agents were already assembled.
People still set the scope, defined the standards, made the decisions that required human authority and accepted the consequences. The agents carried the work among functions between those decisions. The rules gave them no option to tire of another review or decide that everyone was close enough to move on.
I started these experiments wondering how well agents could do expert work. I came away much more interested in what they could do between the experts. A surprising amount of organizational structure exists to keep capable people synchronized.
After 13 hours and 13,475,814 tokens, my artificial department heads reached the end without once asking the question that starts so many human alignment processes:
Can everyone do Thursday at 2?

