A test run of Nyx failed even though its clinical conclusion did not need to change.
We are building Nyx to turn medical evidence into knowledge that another person can review, question and update. In one recent test, it produced an oncology record whose clinical content matched the evidence we had accepted. It also left out three important links: the headline and two statements about the trial results did not point to every fact needed to support their wording.
The facts and the answer were there; what was missing was the path between them. So we stopped the run. The next attempt was allowed to add only the missing links, and we compared the two versions to make sure the medical content, evidence and intended limits had not changed. After the corrected version passed review, a human accepted that exact version into our shared knowledge library. That acceptance alone did not authorize putting it into an app, distributing it, making recommendations from it or using it in patient care.
That is what “right answer” means here. The clinical content matched the accepted evidence and required no medical revision. The record was still incomplete, had not been clinically validated and was not ready for use.
What was missing
As an oncologist, I care whether the medical conclusion holds. As a builder, I also need another person to be able to see exactly why it holds. This run met the first requirement and failed the second.
A shared record has to remain understandable after the person who created it is no longer available to explain it. The support links show which facts sit behind each statement and which statements need another look when a fact changes. A reviewer could have searched through the whole record and reconstructed the missing links. But if every reviewer has to repeat that work, Nyx has not finished its job.
The same problem appears outside medicine. A Microsoft Research study called AgentPex examined AI agents handling airline, retail and telecom customer-service tasks. In one example, an agent correctly cancelled two conflicting reservations but skipped the check that would confirm the cancellations and broke other required steps.
Among 58 test runs from one model that all received a perfect outcome score, AgentPex flagged at least one missed step or broken rule in 48. AgentPex itself used AI to perform the review, and the researchers manually checked only small samples. They say the method should supplement, not replace, human-written test criteria. This was one set of non-medical tests, so it does not tell us how often similar failures occur in healthcare or live systems. It does show why the final result cannot tell the whole story. One test asks whether the task ended correctly. Another asks whether the required work happened along the way. High-stakes work may need both.
The process still has to be worth it
Stopping sound clinical content over missing links can sound like bureaucracy. Sometimes it is. A rule is difficult to defend if it was added after the result, cannot explain the risk it controls or makes people repeat work that was already sound.
The paper raises the same problem from the other direction. If an AI skips a required step and still gets the answer right, it may look efficient rather than unsafe. The important question is why the step was required. A formatting preference and a verification that can catch a silent failure should not carry the same weight. If every small deviation is treated like a serious one, the review process makes it harder to see what actually needs attention.
But here, the rule existed before the run. Another reviewer could see exactly what was missing. The correction was limited to those links, and we compared the versions before and after to confirm that the underlying medical content had not changed.
Four questions before the work moves on
In practice, we ask four different questions. Does the content match the evidence? Can an independent reviewer understand how it was built? Has a human accepted this exact version? What, if anything, has been approved to happen next?
Those questions sound similar, but the decisions are separate. A sound clinical answer does not repair missing support. Passing review does not mean a human has accepted the result. Acceptance into a shared library does not mean an app can use it or that anyone should rely on it in patient care.
Not every AI task needs this much process. A brainstorming draft and a reusable oncology record do not carry the same consequences. The higher the stakes, the more important it is to define what “finished” means before the work begins. In this case, the correction stayed small, the medicine stayed the same and the final decision remained human.

