Seventy pages of an oncology chart enter a generative AI system. One page comes out.
It says: HER2-positive breast cancer. Node-positive at diagnosis. Residual invasive disease after treatment before surgery. Recommendation: trastuzumab deruxtecan, based on DESTINY-Breast05.
The oncologist signs. The patient receives it. She develops interstitial lung disease and is admitted to hospital.
Then someone returns to the 70 pages.
Surgical pathology showed residual cancer in the breast, but no cancer in the lymph nodes. Her cancer was operable. For her, DESTINY-Breast05 required residual node-positive disease after treatment. She would not have entered the trial the system cited. NCCN’s Category 1 preferred recommendation and Canadian consensus recommendations use the same high-risk boundary.
Every clinical fact on that page was true. The system produced the wrong recommendation by leaving out the fact that changed the answer.
Could the oncologist catch it? Yes. She would have to reconstruct the timeline from the 70 pages, reconcile the biopsy with the surgical pathology, retrieve the trial protocol and redo the work the system compressed.
That oncologist will still be called the human in the loop.
A safety mechanism that cannot work
A paper in npj Digital Medicine argues that having a clinician in the workflow does not make oversight meaningful. It proposes four conditions: epistemic capacity, cognitive space, decisional authority and intervention effectiveness. It is a serious attempt to turn a slogan into an operational standard.
Apply that test to a generative AI system making substantive oncology recommendations and the human-in-the-loop collapses. No clinician can meet all four conditions at the same time.
The clinician cannot develop epistemic capacity over the model itself. A model with nine billion parameters contains roughly nine billion learned numerical values. They are not readable clinical rules. Their interactions produce behaviour that changes with the prompt, the surrounding context, the system instructions and the model version.
Open weights allow specialized researchers to investigate those interactions. They do not make the model understandable to a clinician. A closed system restricts access further. The clinician cannot examine the training data, reproduce the evaluations or determine what changed after an update.
Anthropic’s interpretability researchers describe model mechanisms as inscrutable even to developers. Their methods captured only part of the computation and required hours of expert work to examine prompts containing tens of words. That is the current frontier of model inspection, far removed from an oncologist reviewing a complex chart between patients.
Cognitive space creates the next contradiction. Generative AI is introduced because the original task requires more time or attention than the workflow contains. The clinician is then designated as the safety control for the same task. Reliable verification requires reconstructing the work the model was meant to replace.
The authority and intervention conditions fare no better. I can reject an output. I cannot change the model, roll back an update or prevent the same failure from reaching the next patient. Rejecting one summary is clinical judgment. It is not control over the system.
What real human review costs
An oncology study published this year compared trial prescreening by trained research staff with and without help from a language model across 355 charts from patients with lung or colorectal cancer.
Human-plus-AI review improved chart-level accuracy from 71.1 to 76.5 percent. That matters. Average review time was almost identical: 37.8 minutes without AI and 37.4 minutes with it. The reviewers still read every document in each record. The authors describe the work shifting from finding the information to carefully reviewing the AI’s abstractions. The randomized evaluation is worth reading.
The study shows that AI can improve accuracy. It also shows what credible human review cost in that experiment. The full chart review remained. The labour saving did not appear.
That is the oversight contradiction in numbers. Keep the independent review and much of the efficiency disappears. Remove the review and the human loses the basis for detecting omissions. The human cannot rely on the system and independently reproduce its work at the same time.
Transparency cannot manufacture oversight
Canada has opened a consultation on AI transparency, including what people should be told about AI-generated content, system limitations, serious incidents and AI agents. Those are useful questions.
They do not solve this problem.
A model card can tell me which populations were used for validation. A label can describe known limitations. An interface can identify the model and warn me that errors are possible. None of those disclosures lets me inspect nine billion parameters or see a fact that the model omitted.
Transparency can support evaluation by regulators, institutions and technical auditors. Health Canada’s current guidance for regulated machine-learning medical devices asks manufacturers for evidence about architecture, data, validation, subgroup performance, version control and post-market monitoring. That is manufacturer-level evidence. It cannot be transferred to an oncologist through a disclosure form.
Calling the physician the human in the loop gives the appearance of control to someone who cannot inspect the model, cannot verify many of its omissions and cannot alter its behaviour. It also creates a convenient place for responsibility to land when the system fails.
Liability must attach to the output
Canada should adopt a clear rule for generative AI placed into clinical use.
If a company puts the system into an authorized or reasonably foreseeable clinical use, it is responsible for every output produced in that use. When an output materially contributes to patient harm, the company is liable. The patient should not have to prove negligent development, identify a defect hidden within the model or show that the failure had never been disclosed.
Physician review does not reduce that liability. A warning does not transfer it. A signature does not transfer it.
I would call this no-fault clinical-output liability. The European Union’s new Product Liability Directive brings software and AI systems within strict product liability, although an injured person must still establish a defect and a causal link. The directive is an important precedent. A generative-AI provider may argue that a harmful output was expected statistical behaviour rather than a product defect.
The patient should not have to reverse-engineer the commercial stack either. A foundation-model provider, application developer and clinical integrator may all have shaped the output. Where several commercial entities built and supplied the clinical system, they should be jointly liable to the patient and determine their respective shares afterward.
This would allocate the risk to the parties that create the product, control its development, set its price and can insure against its failures. A physician cannot perform any of those functions by clicking approve.
The signature proves very little
The oncologist still has to make a clinical decision. She can question the proposed treatment. But unless she reopens the 70 pages, reconstructs the treatment timeline and checks the trial protocol herself, she cannot see the omitted nodal response that makes the cited evidence inapplicable.
Her signature proves that a doctor was present. It does not prove that the model was supervised.
The human in the loop is a regulatory fiction when it is used to certify safety or shift responsibility. A company that cannot accept liability because its model is unpredictable has explained why that model is not ready for clinical use.

