A clinical note, a clinical decision, and a clinical outcome describe three different things.

Good documentation is an accurate and usable account of the encounter. A good decision applies the available evidence to the patient’s circumstances and values. A good outcome is what everyone hopes follows.

They often align. Medicine also supplies every possible counterexample. A comprehensive note can preserve a weak decision in exquisite detail. A careful decision can disappear into three lines. The best-supported choice can still end badly because biology retains the final vote.

These distinctions matter as artificial intelligence moves into clinical care. A system may improve the record, support the decision, influence the outcome, or do some combination of the three. Each is a separate claim. Each needs evidence at the level where the claim is made.

The claim stops where the evidence stops.

What a large clinical trial showed

A recent pragmatic trial published in Nature Medicine tested an AI clinical decision-support system across 16 primary-care facilities in Kenya. The system was embedded in the electronic medical record. It analyzed the information entered during the encounter and automatically produced diagnostic and treatment guidance. Clinicians could accept, modify, or disregard it.

The investigators randomized 103 clinical officers and enrolled 9,691 patients. The primary analysis included 9,347 encounters. This matters because the study moved beyond a model answering clinical questions in isolation. It tested the system during routine care, with real clinicians retaining responsibility for the decisions.

The clearest positive result appeared in the clinical record.

An expert panel of six family physicians reviewed 2,000 encounters for documentation quality. Compared with usual care, clinicians using the AI-supported system had higher odds of recording:

  • an appropriate diagnosis: adjusted odds ratio 1.74;
  • a comprehensive clinical note: adjusted odds ratio 1.68; and
  • an appropriate treatment plan: adjusted odds ratio 1.71.

The 95 percent confidence intervals excluded 1.0 for all three comparisons, and every P value was below 0.001.

Those are meaningful findings. The diagnosis, the account of the encounter, and the treatment plan form the durable clinical representation of what happened. A complete record allows another clinician to understand the case without reconstructing it from fragments. It makes the reasoning available for review. It preserves why a decision was made when the patient returns weeks or months later.

The effect was also inexpensive at the point of use. The estimated model cost was four US cents per consultation. In a post hoc analysis, antibiotic-related costs were fifteen cents lower per patient in the intervention group. That finding does not establish the full economics of the system, though it provides a useful operational signal.

The primary patient-level outcome was treatment failure within 14 days. This included unresolved symptoms, an unplanned escalation of care, a missed diagnosis, an inappropriate prescription, a serious adverse event, or death. Treatment failure occurred in 2.2 percent of patients in the intervention group and 2.0 percent in the control group. After adjustment for the clustered design, the odds ratio was 0.77, with a 95 percent confidence interval from 0.55 to 1.08 and a P value of 0.13.

The study therefore supports a bounded conclusion. The system improved the quality of what clinicians recorded about the diagnosis and treatment plan. It did not establish a statistically significant improvement in short-term treatment failure. The authors estimated that any true effect on that endpoint was probably modest.

That is a useful result. Better documentation deserves to be assessed on its own terms.

What is a better record worth?

On first reading, I paid more attention to treatment failure. It was the patient-level outcome and the most consequential endpoint in the study. The documentation findings seemed secondary.

That reaction may give the clinical record too little credit.

Imagine opening the chart of a patient who has seen three clinicians over six months. Her treatment was changed twice. One option was considered and rejected. A new problem appeared between visits. The question facing the next clinician is immediate: why is she on this plan?

A thin record supplies dates, diagnoses, and prescriptions. A good record can preserve the reasoning. It shows which evidence was available, which alternatives were considered, what uncertainty remained, and why the plan made sense at the time.

The value often appears later. The next clinician can review the decision without reconstructing it from fragments. A safety review can identify where the process failed. If a complaint or legal proceeding follows, the record can distinguish a reasonable decision with a poor outcome from a decision that lacked a defensible basis.

This is the measured value in the Kenyan trial. Clinicians using the AI-supported system produced records that experts more often judged to contain an appropriate diagnosis, a comprehensive account, and an appropriate treatment plan. The finding is narrower than improved care in general. It is still useful.

This is also a practical design question at Kesis & Sisters. When we organize evidence around a decision, we work to keep the source, assumptions, calculations, and available options visible. The aim is straightforward: another person should be able to follow the reasoning, challenge it, and explain it.

That raises another question. Who else should be able to use the record?

Who is the record for?

Clinical records were built primarily for clinicians and institutions. Patients increasingly read them too.

A comprehensive note can still be difficult for its subject to use. A portal may deliver the exact record while leaving the patient with unfamiliar terminology, disconnected events, and little sense of which findings matter to the next decision.

This week, OpenAI expanded Health in ChatGPT to US users. The system can connect supported medical records, explain a visit note in plain language, summarize changes in laboratory results, and help someone prepare questions for a follow-up appointment.

Those are worthwhile capabilities. They extend the potential value of a good record by making a fragmented clinical history easier to access and understand. A patient who can follow her own timeline, recognize the options her physician discussed, and arrive with better questions has gained something even if her eventual treatment remains unchanged.

That value can be studied directly. Accuracy, comprehension, recall, decisional conflict, and whether the eventual choice reflects the patient’s informed preferences are all measurable. A tool designed to improve understanding should be evaluated through those endpoints. A documentation tool should be evaluated through the quality and usefulness of the record. A system claiming a clinical effect should be evaluated through a patient-level clinical outcome.

The scientific error is category inflation. A benchmark result becomes a claim about clinical reasoning. A better note becomes a claim about a better decision. A better decision becomes a promise of a better outcome. Every transition may be plausible. Each still requires evidence.

Let the result remain its true size

The Kenyan trial is valuable because it gives us a real-world estimate instead of another demonstration. At four cents per encounter, an embedded language model improved the recorded diagnosis, the completeness of the note, and the appropriateness of the documented treatment plan. The trial also placed boundaries around the likely short-term clinical effect.

That is what credible progress looks like in a field moving this quickly. Name the intended value. Choose an endpoint that measures it. Report the result at the level the study can support.

For a patient, good documentation, a good decision, and a good outcome are three different goods. She deserves an accurate account of her care. She deserves a decision she can understand and defend. She deserves our best effort toward the outcome she hopes for.

We can build systems that contribute to each. Science should tell us which one improved.

The claim stops where the evidence stops.


Dr. Henry Conter is a Medical Oncologist and Hematologist at William Osler Health System and the founder of Kesis & Sisters. He trained in Medical Oncology at MD Anderson Cancer Center and spent six years at Hoffmann-La Roche in progressively senior roles spanning oncology clinical development, portfolio strategy, and medical and regulatory affairs, across both national and global functions.