Overall survival has occupied the top of the oncology endpoint hierarchy for most of my career. When I worked with Canada’s Drug Agency, reviewing cancer medicines through pCODR, I treated OS as the hard outcome. PFS was a surrogate, and response rate sat another step away. My friend Bishal Gyawali has spent years arguing that cancer medicines should be judged by outcomes that matter to patients. I agree with him.
A new JAMA research letter from Brian Shkabari and colleagues, including Bishal, examined 385 adult solid-tumour indications authorised by the FDA between 2006 and 2025. The proportion supported by OS fell from 40.0 percent in 2006–2010 to 18.7 percent in 2021–2025, while reliance on surrogate endpoints rose from 60.0 to 81.3 percent. Among regular approvals, use of surrogate endpoints increased from 53.8 to 74.8 percent. The authors see a concerning movement away from survival and toward PFS and response rate. A few years ago, I would have read their findings the same way.
The paper instead forced me to look inside OS itself. For a patient who progresses before death, OS combines two intervals with different evidentiary properties: time from randomisation to progression, followed by survival after progression. During the first interval, the treatment and comparator, dosing, assessments and rules for discontinuation are defined by protocol. After progression, the trial may continue to count deaths, although the original randomisation usually no longer defines the care that helps determine when those deaths occur. The issue is therefore structural. As post-progression survival grows, a larger share of OS arises during a phase with weaker treatment attribution and lower transportability than the phase the trial was designed to control. I have come to think that PFS is often the stronger endpoint for treatment-specific efficacy, while OS carries the separate and essential role emphasised in the FDA’s recent draft guidance: detecting a signal of fatal harm.
One endpoint, two phases of evidence
OS is defined as the time from randomisation until death from any cause. The event is objective and its importance to patients is self-evident. Neither property, however, tells us how specifically the treatment contrast estimates the effect of the indexed drug. For patients who progress before death, OS can be decomposed into the pre-progression interval and post-progression survival. The first is observed largely within a randomised, protocol-defined treatment comparison. The second is observed during survival follow-up, after the patient’s therapeutic course has branched. Randomisation still balances prognosis at baseline, and an intention-to-treat OS comparison remains internally valid for the effect of assigning one initial treatment strategy rather than another in the trial as conducted. It is a broad estimand, encompassing consequences mediated through subsequent care as well as the indexed treatment’s efficacy and toxicity. That breadth is clinically meaningful, while offering less specificity about what the first treatment contributed.
There is a second statistical problem inside that branch. The patients who progress are no longer a random subset of the population enrolled in the trial. Progression is a post-randomisation event influenced by both the assigned treatment and the patient’s prognosis. If the experimental treatment delays progression, a patient who nevertheless progresses on it may have a different underlying risk profile from a patient who progresses on control. Restricting a comparison to progressors therefore conditions on a common consequence of treatment and prognosis. Statisticians call this collider or selection bias. As Xabier García-Albéniz, Joan Maurel and Miguel Hernán have shown, the treatment groups can cease to be comparable within that selected subgroup even though randomisation made them comparable at baseline. This selection problem does not invalidate the full intention-to-treat OS comparison, which includes every patient according to randomised assignment. It arises when we isolate those who progressed, or interpret survival after progression as though it were a second randomised comparison.
The problem then becomes recursive. The first treatment can change tumour biology, toxicity, performance status, elapsed time and the options available at progression. Clinicians observe an incomplete version of those changes, form a judgment about what has happened and choose the next treatment. That treatment changes the patient’s next state and the next decision. In causal inference, this is time-dependent confounding affected by prior treatment, often called treatment-confounder feedback: factors affected by earlier treatment both influence later treatment and predict survival. Its direction cannot be assumed. The first treatment could sensitise or select for resistance, preserve or diminish fitness, and open or close later options. A clinician’s interpretation could steer the next choice in either arm. Only the initial treatment was randomised. The intention-to-treat OS estimate consequently belongs to the adaptive pathways that began with that assignment. It remains valid for comparing those initial strategies as they unfolded in the trial, while becoming less specific to the indexed drug.
External validity is usually discussed in terms of whether trial participants resemble patients in practice. OS introduces another transportability question: whether care after progression resembles the care available where, and when, the result will be used. Eligibility criteria limit generalisability during both periods, yet the pre-progression treatment contrast remains defined and reproducible. Another centre can identify the regimen, comparator, dosing and assessment schedule it is being asked to apply. Post-progression care is a distribution of pathways rather than a single reproducible exposure. Transporting a PFS result requires a comparable patient, treatment contrast and assessment process. Transporting an OS result also requires the new setting to reproduce the distribution of later therapies and decisions in both randomised groups. That distribution varies across jurisdictions, institutions and calendar time. Post-progression care may reflect clinical practice, but its realism is local. This makes the post-progression component of OS substantially less transportable than the protocol-defined component that precedes it.
Crossover, subsequent therapies, supportive care, access, clinician and patient choice, and competing illness are the mechanisms through which this second phase varies. Their importance differs across diseases, and a trial may protocolise or adjust for some of them. An adjustment for crossover cannot standardise a therapy approved midway through follow-up. Equal access within a trial cannot reproduce access in another health system. A table of subsequent drugs cannot capture every decision that shaped survival. The common source of heterogeneity remains: after progression, patients enter multiple care pathways that the original randomisation did not assign. The ICH E9(R1) estimand framework calls treatment switching, discontinuation and additional therapy intercurrent events because they change the clinical question attached to the treatment effect. Following patients through those events yields a legitimate estimate of an initial strategy within a particular downstream care ecosystem. It answers a broader question than whether the indexed drug controlled the cancer.
The evidence base often leaves that downstream ecosystem only partly visible. In a cross-sectional study of 275 published randomised trials and 77 trials supporting FDA approvals from 2018 to 2020, assessable post-progression treatment data were available for 36.4 percent of the published trials and 48.1 percent of the registration trials. A separate 2024 study of 334 phase III oncology trials enrolling 265,310 patients found that post-progression therapies were reported in 47 percent and accounted for analytically in 12 percent. PFS and OS interpretations were discordant in 32 percent of the trials; among the 42 reporting crossover rates, greater crossover was associated with greater odds of discordance. These studies do not settle which endpoint should lead in every disease. They show that the phase contributing an increasing share of OS is influential, incompletely characterised and difficult to reproduce outside the original trial.
The statistical cost of that extension has been quantified. In a 2009 simulation by Kristine Broglio and Donald Berry, a trial was designed around an improvement in median PFS from six to nine months. Detecting that PFS difference required 280 patients. Detecting the same carried-forward effect in OS required 350 patients when median survival after progression was two months and 2,440 patients when it was 24 months. With a PFS result at P = .001, the probability of also finding a statistically significant OS result fell from more than 90 percent to less than 20 percent across those two post-progression survival scenarios. The simulated treatment effect remained fixed; variability after progression overwhelmed the trial’s ability to identify it through OS.
Clinical data show the same pattern. A meta-analysis of 26 first-line ovarian cancer trials involving 24,870 patients found that the trial-level relationship between PFS and OS weakened as treatment evolved. The reported r² fell from 0.66 in the pre-platinum and paclitaxel era to 0.30 in trials of biological therapies. The authors linked that erosion to longer post-progression survival and greater access to effective salvage treatment. These findings help explain why a treatment can produce a clear PFS effect and an inconclusive OS result even when the original benefit carries forward.
PFS narrows the question to the period before progression or death, usually before post-progression treatment changes the therapeutic exposure. DFS does the same in curative-intent care by asking whether an adjuvant treatment delays recurrence or death. Both endpoints have sources of error, including scan timing, informative censoring, assessment bias and inconsistent definitions. Prespecified schedules, randomisation and blinded review make much of that measurement error visible and testable. Their advantage is causal proximity and reproducibility: they remain close to the treatment decision the trial was designed to test. OS retains greater efficacy value when natural history is short, later treatment is limited and death occurs close to randomisation. The FDA uses metastatic pancreatic cancer as an example. The argument for PFS becomes stronger as post-progression survival grows, subsequent pathways multiply and the downstream care observed in the trial becomes less likely to travel with its result.
Authorisation and sequence are different decisions
Calling PFS a direct efficacy endpoint still leaves a second decision. Market authorisation asks whether a drug has shown sufficient efficacy, with an acceptable safety profile, to be available in a defined setting. A mandatory standard-of-care claim asks whether patients should receive it at that point rather than retain it for later. PFS can answer the first question by showing that treatment delays progression. A PFS result still has to earn that role. The magnitude and durability of the effect, the reliability of progression assessment, toxicity, symptoms, quality of life and available alternatives determine whether it supports authorisation. Answering the second question requires evidence about sequence. The distinction matters whenever the same active drug can be given after progression.
Alyson Haslam and Vinay Prasad offered a useful framework for crossover in 2018. They argued that crossover is desirable when a drug has already shown benefit in a later line and a trial is attempting to move it earlier. Appropriate crossover then gives the control group access at progression, and the trial asks whether immediate use is better than deferred use. In the five-year analysis of KEYNOTE-024, for example, PD-1 therapy was already established after first-line treatment, and 66.0 percent of patients initially assigned to chemotherapy received subsequent anti-PD-1 or PD-L1 therapy. Median OS nevertheless remained 26.3 months with first-line pembrolizumab and 13.4 months with chemotherapy (HR 0.62; 95% CI 0.48–0.81). Their framework treats crossover differently when the fundamental efficacy of an experimental drug remains unknown. Giving that drug to the control group after progression can obscure its mortality effect and may delay an effective alternative. Crossover therefore changes the estimand. Its value depends on whether the trial is testing basic efficacy or treatment sequence. In a sequencing trial, OS’s pathway-level nature becomes useful. PFS measures what immediate exposure did before progression; OS asks whether moving the same active drug earlier changed survival across the whole sequence.
Seen this way, the FDA’s regulatory approach and Bishal’s evidentiary concern are compatible. A PFS improvement can support authorisation in the earlier setting because it demonstrates that the drug delays progression when used there. Appropriate crossover preserves a separate question: whether moving an already effective later-line drug forward improves survival across the full treatment pathway. If the OS analysis does not demonstrate an advantage for earlier use, the authorisation can remain justified while the sequencing claim remains unsettled. The trial has established that the drug works in the earlier setting; it has not established that every eligible patient must receive it there. The two sequences may still differ in toxicity, quality of life, cost, treatment-free time and the chance of reaching later therapy. Equivalence would require its own design and evidence. Market authorisation and a mandatory standard-of-care claim require different evidence.
Why the FDA is separating efficacy from safety
The FDA’s final guidance on oncology trial endpoints already recognises that PFS, DFS and EFS can represent direct clinical benefit. The determination depends on disease setting, effect size, available therapy and the clinical consequences of delaying progression. Preventing a new brain or spinal lesion, postponing recurrence after curative-intent treatment or delaying a more toxic therapy can matter to a patient in their own right.
The agency’s August 2025 draft guidance on assessing OS still calls OS the gold standard and says it should be prioritised as the primary endpoint when feasible. I take that language seriously. I also take the agency’s decisions seriously. As Shkabari and colleagues show, by 2021–2025, 81.3 percent of solid-tumour indications were supported by surrogate endpoints, including 74.8 percent of regular approvals. I read that record as evidence that, whatever OS’s formal place in the endpoint hierarchy, a demonstrated OS benefit is increasingly unnecessary as the efficacy basis for market authorisation. The agency is repeatedly accepting other endpoints as sufficient evidence of efficacy.
Survival follow-up remains essential, with a different task. The draft describes OS as both an efficacy and safety endpoint, while focusing primarily on trials in which response, PFS or EFS provides the main efficacy result. In those trials, FDA recommends a prespecified OS safety analysis designed to assess potential harm. It also recommends that all randomised oncology trials collect survival data and plan analyses for harm even when OS has no formal place in the efficacy testing hierarchy.
Observed discordance between endpoints explains the need. FDA authors reviewed six randomised trials in lymphoid cancers and three trials in recurrent ovarian cancer in which PFS improved while later OS results suggested possible harm. They also identified immunotherapy trials with OS benefits that conventional PFS or response criteria did not capture. Their paper, “Irreconcilable Differences: The Divorce Between Response Rates, Progression-Free Survival, and Overall Survival”, argues for rigorous survival follow-up because early antitumour efficacy can be outweighed by toxicity.
BELLINI makes that safety role concrete. In the final analysis of the randomised phase III trial of venetoclax added to bortezomib and dexamethasone for relapsed or refractory multiple myeloma, median PFS was 23.4 months with venetoclax and 11.4 months with placebo (HR 0.58; 95% CI 0.43–0.78). The OS point estimate favoured placebo (HR 1.19; 95% CI 0.80–1.77), and four treatment-related deaths occurred with venetoclax compared with none in the control group. The trial authors concluded that venetoclax should be avoided in the general relapsed or refractory myeloma population. The PFS result demonstrated disease control. OS exposed the benefit-risk problem that PFS could not capture.
KEYNOTE-024 supplies the other direction. PFS established pembrolizumab’s treatment effect before progression; positive OS despite substantial crossover added evidence that using it earlier improved the treatment pathway. The interpretation is asymmetric. A credible harmful OS signal can overturn or narrow an otherwise favourable PFS result. Positive OS strengthens the case for the treatment and for its tested position in the sequence. An inconclusive OS result does neither. It leaves the PFS efficacy result intact while leaving the sequencing question open. OS is therefore weaker for treatment-specific efficacy, yet powerful as a broad warning of net harm and as confirmatory evidence about the whole treatment strategy.
This division of labour follows the strengths of each endpoint. PFS, DFS and EFS estimate whether the assigned treatment changed the course of the cancer during the relevant treatment window. OS integrates fatal toxicity, complications, discontinuation and downstream consequences into a broad mortality signal. That breadth gives OS its safety role: it can reveal whether early antitumour activity is offset by fatal harm. The same breadth weakens its specificity, and often its statistical power, as an efficacy measure of the indexed drug. In much of modern oncology, PFS has become the primary efficacy endpoint and OS the integrated safety endpoint. OS does not measure nonfatal toxicity or quality of life, and an inconclusive OS result cannot prove the absence of harm. In safety assessment, its role is to detect unacceptable mortality risk alongside adverse-event and patient-reported outcome data.
Response rate is now the surrogate
Once PFS is understood as the direct efficacy endpoint, response rate assumes the surrogate role. ORR measures whether a drug produced a predefined amount of tumour shrinkage. The effect occurs early and is especially useful in a single-arm study because a substantial response in a refractory cancer is more readily attributable to the drug than a time-to-event outcome without a control group. Depth and duration of response then help estimate whether that activity is likely to become durable disease control, which a randomised PFS comparison later tests directly.
The relationship varies by disease and treatment mechanism. Cytostatic drugs can extend PFS with few formal responses, while a short-lived response can raise ORR without meaningfully delaying progression. Immunotherapy introduces further differences between radiographic response and longer-term outcome. A 2024 meta-analysis of 675 phase III trials found an R² of 0.33 between treatment effects on ORR and PFS across pooled tumour and treatment types. That modest overall relationship supports the FDA’s contextual approach: disease, effect size, depth, duration and available therapy determine what a response result can reasonably predict.
The resulting sequence is clinically coherent. ORR provides an early estimate of antitumour activity. PFS, DFS or EFS determines whether the activity changed the course of disease. OS follows long enough to identify a mortality signal after efficacy, toxicity and subsequent care have all exerted their effects. Each endpoint answers a different question at a different point in the pathway.
OS is the endpoint of the pathway
Cancer treatment is a sequence. A clinician chooses a therapy, observes response, changes course at progression and chooses again. By the time OS is observed, several later decisions may have shaped it. In machine learning this is a credit-assignment problem: a delayed outcome follows multiple intervening actions, making the contribution of the first action difficult to isolate. OS is the endpoint of the treatment pathway. PFS sits closer to the current treatment decision.
That distinction matters in our clinical decision-support work at Kesis & Sisters. We often use PFS or DFS because the outcome aligns with the decision being modelled. Toxicity, quality of life and the OS safety signal remain visible as separate constraints. This structure helps a clinician distinguish the effect attributable to the current treatment from outcomes produced by the larger sequence of care.
For years, I asked whether PFS was a good surrogate for OS. I now ask which endpoint best estimates what the assigned treatment did to the cancer, and which endpoint best detects fatal harm. In many contemporary settings, PFS answers the efficacy question with less downstream heterogeneity. OS answers the safety question by integrating everything that follows.
Bishal and his colleagues have documented a major change in the evidence supporting cancer medicines. Their analysis asks whether regulators have lowered the evidentiary bar. I see endpoints being assigned different jobs as survival lengthens and treatment pathways become more complex. The positions can coexist. PFS can establish disease control sufficient for authorisation; OS and sequencing evidence determine how far that result should govern the treatment pathway. Their paper changed my view, although perhaps it narrowed our disagreement more than it placed us on opposite sides. Given how often I have agreed with Bishal on this subject, I am interested to hear where he thinks this boundary belongs.
A patient wants to live longer, and a trial needs to tell us what the assigned treatment contributed to that goal. Modern oncology needs an efficacy endpoint close enough to the treatment to identify its effect and a safety endpoint broad enough to detect its harm. It also needs to separate evidence that a drug works from evidence that it must be used at a particular point in the sequence. PFS can support the first claim and therefore often provides the efficacy basis for market authorisation. Authorisation makes the drug available. Where and when it belongs in an individual patient’s treatment sequence remains a clinical question, informed by comparative sequencing evidence and decided at the bedside by the patient and clinician together.

