When someone shows me a medical prediction model, I like to start with who it is meant to help. Then I look at whether the people it learned from, and the people it was tested on, resemble those patients. That gives me somewhere to begin when we get to the question of how well it works.

It is much the same conversation I have with myself when reading a clinical trial. There may be an encouraging result, and I want it to be useful. Somewhere in the back of my mind is the person who might ask me, “Does that mean it would help me?”

To work towards an answer, I look at who took part, how the information was collected, what was measured and what is missing. These are familiar questions. I find them just as helpful when the paper is about artificial intelligence.

Who are the people behind the result?

Say a cancer trial finds that a treatment helps people live longer. It has my interest. But if everyone in the trial was receiving their first treatment, and the person I am thinking about has already had several, I pause. How much of that result can I reasonably bring into our conversation?

Different patients do not make a trial useless. I try to work out whether the differences matter. That is easier when the authors give me a clear picture of the people they studied. It becomes much harder if I cannot find out where the patients came from or how their diagnoses were established. I am left trying to make a connection without knowing what is on the other end.

That is what caught my attention in a study by Alexander Gibson and colleagues in BMC Medicine. They examined two Kaggle datasets concerning stroke and diabetes and could not verify their authenticity. Both scored zero on nine selected reporting items about the origins of the data. Yet 125 published prediction studies had used them. The stroke dataset’s uploader had described it as suitable for education only, excluding research and commercial use. And in April, Scientific Reports retracted a stroke-prediction paper after its authors could not provide further information about the provenance or accuracy of the dataset used to train and assess its models. The editors lost confidence in the findings; the authors disagreed.

I can understand the appeal of an accessible dataset and a promising new method. But when I come back to that patient’s question, I still need to know whose experience the answer rests on.

How many different looks at the same patients?

A stack of papers can feel reassuring. I find it useful to look underneath the stack and ask how many separate groups of patients are actually there.

An initial trial report, a longer follow-up and several subgroup analyses may all help me understand a treatment. They still come from one trial. Linking those reports is a familiar part of a systematic review, described in the Cochrane Handbook. I bring the same thought to AI: ten teams working with the same patient records have given me ten analyses, but not experience in ten different patient populations.

There is another wrinkle. In Silberzahn and colleagues’ soccer study, 29 teams analysed the same dataset to ask whether darker-skinned players were more likely to receive red cards. Twenty teams found a statistically significant association, all pointing towards more red cards for darker-skinned players. Nine did not reach statistical significance. The analytical choices affected how large the association appeared and how confidently it could be distinguished from no association. I wonder how confident I might have felt if I had happened to read just one of those analyses.

Patel, Burford and Ioannidis explored this in health data, testing 8,192 combinations of adjustment variables for each of 417 associations with mortality. For 31% of those associations, the estimates spanned opposite directions. Depending on the adjustments, the same factor could appear associated with a higher or a lower risk of death.

What helps me here is being able to see the alternatives. Simonsohn, Simmons and Nelson’s specification-curve analysis offers a way to compare results across reasonable, statistically valid analytical choices. I find that a generous thing to show a reader: here is what we found, and here is how much the answer moves when we make different choices. It helps me see where my confidence is well placed and where I still have questions.

Kerrington Powell and Vinay Prasad make a related point about portfolios of clinical trials. With enough tests, some can be positive by chance, even when those tests sit in separate studies. When a result looks promising, I like to know what else was tried and how it turned out.

That leaves me with a useful question for medical AI too: how does this result hold up beyond the analysis that first made it look promising? Seeing it work in new patients would give me a different kind of reassurance from seeing another paper about the same records.

And what might this mean for you?

Eventually, I come back to the conversation that makes all this reading worthwhile. Suppose a model suggests that someone is at high risk of a complication. I can imagine their next question: “Is there anything we can do about it?” That takes us from the prediction to the available choices. I would look at the evidence for their benefits and harms, then talk through what those choices would involve. A treatment that looks worthwhile on paper may feel quite different to someone weighing its burden against the things they hope to keep doing.

A risk estimate can help that conversation along. Finding out whether acting on it helps takes further evidence. Finding out what matters to this person takes listening. I find it reassuring that there is already guidance for much of this work. TRIPOD+AI sets out reporting requirements for prediction models, including their intended population, data sources, participants and outcomes. For the technical questions, I may need help from someone with different expertise. I can still begin with the clinical questions I know how to ask.

That is why I keep coming back to evidence-based medicine. It gives me a place to start with something new. I look at who was studied, how dependable the finding is and whether it helps with the decision at hand. Then I try to understand what the person in front of me hopes to gain and what they are willing to accept. I do not feel a need to reinvent that approach for artificial intelligence. I want to carry it with me. The technology may be new; the wish to give someone a useful, honest answer is very familiar.