When I read a randomized trial, I usually begin with Table 1. Before looking at the outcome, I want to know who the intervention was designed for and who was actually studied. Age, disease stage and prior treatment give me the first outline of the population to which the result belongs.

Doctors then package what the trial taught us into patient types we can recognize later. We remember the newly diagnosed patient with a particular stage and biomarker, or the heavily treated patient who remains fit enough for another therapy. These working categories help us retrieve the relevant evidence and notice when we have begun to extrapolate beyond it.

Reading a recent Nature Medicine study of explainable artificial intelligence in dermatology, I tried to do the same thing with the people using the system. Participants were described as members of the public, medical students or primary care physicians. Those labels left an important characteristic unstated. I wanted to know how each group performed on the task before the system appeared.

Part of the design preserved each participant’s unaided decision before showing the system’s advice. That choice shows how the performance of an artificial intelligence system is entangled with the person using it. Intended users are part of the conditions under which benefit and risk are produced.

Before the system appeared

Participants included 623 members of the public, 153 primary care physicians and, in a separate comparison, 320 medical students. Dermatologists did not take part in the diagnostic experiment. Three dermatologists separately rated the system’s explanations for correctness and informativeness, so the study provides no measure of how they would have performed with the tool.

Members of the public distinguished melanoma from benign naevi, commonly called moles. Physicians and students faced a harder, open-ended differential diagnosis across 34 skin diseases. Each participant assessed 12 images with one of four forms of assistance, ranging from a confidence score to a natural-language explanation.

The investigators called one part of the design Human-First. Participants recorded an initial diagnosis and their confidence, then saw the system’s output and made a final decision. The investigators could observe unaided performance and the response to both accurate and inaccurate advice.

Among the 96 primary care physicians assigned to Human-First, the correct diagnosis was their first choice in 11.5 percent of cases without assistance. The study calls this top-1 accuracy. It appeared somewhere among their three proposed diagnoses in 16.1 percent of cases, the top-3 accuracy. After the system’s advice, those measures rose by 21.5 and 43.5 percentage points.

Those starting values are easy to misread. Participants assessed images without a history, physical examination or the other information available in an ordinary consultation. The numbers describe performance within this experiment rather than the physicians’ general diagnostic ability.

Incorrect recommendations were especially informative. Among the 90 members of the public assigned to Human-First with a natural-language explanation, accuracy rose by 13.4 percentage points when the recommendation was correct and fell by 21.1 points when it was wrong. The same explanation that made good advice useful also made bad advice more persuasive.

The public and physician results should not be compared directly because the tasks differed. The cleaner comparison was between primary care physicians and medical students completing the same differential diagnosis. After adjustment, medical students were 6.7 percentage points more likely to be classified as deferential to the system. Participants reporting greater knowledge of skin disease were also less deferential.

The study classified participants as deferential when they were correct only when the system was correct and wrong whenever it was wrong. Among the 96 primary care physicians in Human-First, 79 met that definition and 17 did not. The deferential group had started with top-1 accuracy about 19 percentage points lower. The groups were defined after the responses were observed, expertise was not randomized and each participant encountered only a few incorrect predictions. This is an exploratory finding. It nevertheless raises a question worth carrying into future studies: could the people with the most room to benefit from correct assistance also have the most to lose when the assistance is wrong?

What a job title leaves out

Guidance for reports of randomized trials involving artificial intelligence and its companion for trial protocols already recognize that the user matters. They ask researchers to describe intended users, required expertise, training and the way people interact with the system. Guidance for early clinical evaluation of decision-support systems goes further by treating users as study participants and asking about baseline characteristics, human-computer agreement and learning curves.

These guidelines make the intended user visible. Their usual descriptors, such as specialty, seniority and years of experience, remain proxies for the capability that matters to a particular task. A medical oncologist with years of practice may be a novice at classifying dermatological images. A primary care physician with a particular interest in dermatology may have far more relevant experience.

A clinician baseline table could narrow that gap without becoming a league table of doctors. Any capability measure belongs to one task. Its purpose is to describe the starting conditions and identify which of them might change the effect of the system’s assistance.

When the evidence meets the market

Commercial positioning often describes a product simply as artificial intelligence for clinicians. The study made that description feel underspecified. Its evidence applies to particular users completing a defined task within a particular workflow.

Imagine offering the same software to medical students, primary care physicians and dermatologists. These look like three customer segments for one product. Each group could start with different task-specific capability and respond differently to both the recommendation and its explanation.

Value may change with the user. Someone with less experience on the task might gain more accuracy while needing additional training or a clearer escalation path. An expert might value consistency or time more than an accuracy gain. These remain product hypotheses; the dermatology study did not test them.

In its August 2026 human-factors guidance, the United States Food and Drug Administration approaches the problem from usability. For human-factors validation, it treats user populations as distinct when they perform different tasks or differ in knowledge, experience or expertise in ways that could affect their interaction with the interface and their potential for use error. Clinical effectiveness requires its own evidence, grounded in an equally clear account of who used the intervention.

A commercial claim needs the same specificity. Evidence from primary care physicians assessing image-only dermatology cases does not automatically describe specialists, complete clinical encounters or months of repeated use. A new user population may require a different product design and a different measure of value.

Reading the result differently

This experiment used 12 image-only cases per participant, deliberately included incorrect system recommendations and measured no patient outcomes. Routine clinical performance remains an open question. Its distinctive contribution was preserving each participant’s unaided decision.

To make that starting point visible, I would separate the clinician baseline table from the outcomes observed after the system was introduced. The baseline could include:

  • Clinical context: role, specialty, training stage and practice setting.
  • Experience with the task: relevant case volume, recent exposure and formal training.
  • Unaided performance: accuracy on a separate set of cases before the system is introduced.
  • Confidence calibration: whether confidence rises and falls with actual accuracy.
  • Prior exposure to decision support: previous use of similar tools and any training provided before evaluation.

Response to incorrect advice and improvement with repeated use also matter, although they occur after the system has been introduced. They belong with the study outcomes rather than in the baseline table.

The table itself remains descriptive. Imagine that the system improves accuracy by 20 percentage points among clinicians with lower baseline performance and by five points among clinicians with higher baseline performance. Showing that each subgroup changed, or that one change was statistically significant and the other was not, would not establish that the effects were different. A test of interaction compares the two effect estimates directly. This is the relevance of a statistical note on interaction: it provides the general trial method for testing whether a baseline characteristic modifies an intervention’s effect. It is not specific to artificial intelligence.

Many experiments of this kind ask each clinician to assess several cases and show each case to several clinicians. A 2023 clinical vignette study had this design and used a cross-classified model to account for variation between both clinicians and cases. That study did not test task-specific clinician capability as an effect modifier. Its relevance here is narrower: it shows how to handle those two sources of variation.

A future trial could use a similar model to ask whether the change in accuracy after assistance depends on what the clinician could do beforehand. The capability measure should be chosen before the trial, and the trial should be large enough to test the interaction. Rather than divide clinicians into arbitrary high and low groups, the analysis could keep capability continuous. The model could then show the expected percentage-point benefit or harm across the observed range, with uncertainty around the curve.

Medical education offers one possible way to improve the baseline measure. In a 30-item clinical-reasoning examination, item response theory was used to model both learner ability and question difficulty. Applied to this setting, the same approach could help estimate task-specific capability while allowing some cases to be harder than others. It would require many more separate unaided cases than the 12 used here, along with evidence that the measure was reliable for this task.

Together, the table and the interaction model would give the result a population. The table shows who used the system and where they started. The model tests whether that starting point changed what happened next.