When someone asks how a cancer treatment is going, the first answer is rarely a survival estimate. We say the cancer has disappeared on the scans, shrunk, remained stable, or grown. We may add that the growth is slight or rapid, that one lesion is behaving differently from the others, or that the scan looks good even though the person does not feel well. Long before any of those observations become a clinical-trial endpoint, they are our shared language for whether treatment is controlling the disease.

In “Overall Survival Is the Weaker Endpoint”, I wrote about a JAMA research letter from Brian Shkabari and colleagues, including my friend Bishal Gyawali. It examined 385 adult solid-tumour indications authorised by the US Food and Drug Administration between 2006 and 2025. In the most recent five-year period, 81.3 percent were supported by surrogate endpoints and 18.7 percent by overall survival. The paper reports surrogate endpoints collectively, including response rate as well as progression-free survival, but the direction is clear. FDA solid-tumour approvals have moved away from requiring a demonstrated improvement in overall survival.

That movement is difficult to explain by statistical surrogacy alone. A meta-analysis of 675 phase III trials involving 350,112 patients found that treatment effects on PFS explained 38 percent of the variation in treatment effects on OS across the pooled trials. The relationship varied substantially by tumour and treatment type. PFS is not a dependable universal predictor of how much longer a treatment will help people live.

Yet progression has become one of the central organizing events in oncology trials, regulatory decisions, health technology assessment, clinical practice, and the systems we build to support those decisions. I do not think that happened because everyone forgot the limitations of surrogacy. I think PFS was adopted so readily because it formalized a model of cancer that patients, clinicians, and almost anyone hearing the words already understood. That fit also sets the standard by which each PFS analysis should be judged.

We had chosen the endpoint before we named it.

Cancer control already has an intuitive order

The word surrogate puts PFS in a subordinate position. It invites one question: how accurately does PFS predict OS? That is an important question when PFS is being used to make a claim about survival. It is not the only reason to measure PFS. An endpoint can also matter because it describes the state treatment is trying to preserve.

The surrogacy framing can produce a bedside sentence that no oncologist could deliver with a straight face: “Your cancer has shrunk by 25 percent! Technically, the trial calls that stable disease. But progression-free survival does not reliably predict overall survival, so perhaps we should feel ambiguous about the result.”

That is obviously absurd. The cancer has shrunk. If that shrinkage holds, it is a good thing. Whether it shrank by 25 percent or 30 percent changes the trial label, not what the patient and oncologist can see.

PFS is an imperfect measure of cancer control, perhaps the worst one except for all the others we have tried. It compresses very different states into one category. But it captures something real: the time during which the cancer is gone, smaller, or at least not meaningfully worse. It does not tell us how long the patient will live. It tells us how long that control lasts.

On the axis of tumour control, the possible states have an intuitive order. No detectable disease on the relevant assessment is better than substantial shrinkage. Substantial shrinkage is generally better than modest shrinkage. Modest shrinkage is better than no change, and no change is better than growth. Slow growth is preferable to rapid growth. New lesions, organ compromise, or accelerating disease are worse again. Toxicity, symptoms, quality of life, treatment burden, and patient preference can change the overall judgment, sometimes decisively, but they do not make the direction of tumour control mysterious.

This is close to how an oncology visit proceeds. We compare the current scan with the prior one. We examine the direction and pace of change. We ask how the person is feeling, what the treatment is costing them physically, and what other options remain. Then we decide whether the current treatment is still doing enough to continue.

PFS adds time to this picture. It counts how long a patient remains alive without the cancer progressing. During that time, the cancer may have disappeared, shrunk substantially, shrunk a little, remained unchanged, or even grown without yet meeting the criteria for progression. The PFS curve puts all of those states together, while patients and clinicians continue to see the differences among them.

Stable disease is a crowded category

The Response Evaluation Criteria in Solid Tumours, or RECIST 1.1, give trials a standardised way to classify what is seen on imaging. A complete response means the disappearance of all target lesions under the criteria. It does not prove that every cancer cell has been eliminated. A partial response generally requires at least a 30 percent decrease in the sum of target-lesion diameters from baseline. Progressive disease generally requires at least a 20 percent increase from the smallest sum recorded during the study, together with an absolute increase of at least 5 mm, or the appearance of new lesions. Stable disease is what remains: neither enough shrinkage to qualify for partial response nor enough increase to qualify for progression.

That residual definition, and the wider progression-free state in which it sits, contain several trajectories that people understand differently. The exact RECIST label at a later assessment can depend on the prior response and the smallest measurement recorded. Without proposing new categories, it is useful to divide the progression-free space conceptually:

  • modest shrinkage that does not reach the partial-response threshold;
  • measurements that are approximately unchanged; and
  • genuine growth that has not reached the progression threshold.

The categories could be placed on a simple continuum: no detectable target disease, substantial shrinkage, modest shrinkage, approximately unchanged disease, subthreshold growth, and formal progression. Everything before the last category may still be progression-free. That does not make modest shrinkage and slow growth equivalent. It means neither has crossed the boundary chosen to mark loss of control.

Direction and speed add another layer. A 10 percent increase over six weeks and a 10 percent increase over six months may receive the same categorical label at an assessment, although they do not imply the same disease kinetics. A tumour can also be growing more slowly on treatment than it would have grown without treatment. In that case, the disease state has worsened while the treatment may still be having an effect. State and treatment effect are related, but they are not identical.

RECIST was not designed to preserve every part of that information in one label. It trades biological detail for a rule that different investigators can apply across centres and time. That trade is not careless. Imaging measurements contain noise. In a study in which people with non-small-cell lung cancer underwent repeat CT scans within 15 minutes, 84 percent of repeated measurements were within 10 percent of one another, but apparent changes ranged from 23 percent shrinkage to 31 percent growth. Three percent of measurements would have met the relative RECIST threshold for progression on immediate repeat imaging. The absolute 5 mm requirement and the distance between stability and progression help prevent every small fluctuation from becoming a treatment-changing event.

The categories still discard information. In a small early-phase study of 76 patients, tumour growth accelerated relative to the pretreatment period in 38 percent of patients classified as non-progressive at 12 weeks. Conversely, growth slowed in 53 percent of those classified as progressive. The study was not a definitive test of RECIST, but it makes the limitation plain: a threshold records whether a line has been crossed, not the complete trajectory that led there.

That limitation does not make the line useless. Blood pressure, kidney function, cardiac ejection fraction, and laboratory toxicity grades all turn continuous measurements into categories when a decision requires one. No serious clinician believes that 19 percent tumour growth is biologically favourable while 20 percent is suddenly disastrous. The threshold makes a continuous process adjudicable. Its value depends on what the boundary is being asked to do.

Health technology assessment uses the same map

The same mental model appears in the mathematics used to evaluate cancer treatments. A common oncology cost-effectiveness model is the three-state partitioned survival model. At any time, it divides the trial population into people who are progression-free, people who are alive with progressed disease, and people who have died.

The arithmetic is simple:

  • the proportion in the progression-free state comes from the PFS curve;
  • the proportion alive with progressed disease is the OS curve minus the PFS curve; and
  • the proportion who have died is one minus the OS curve.

OS tells the model who is alive. PFS supplies the distinction between two very different conditions among those people who are alive. In a recent endometrial-cancer appraisal, NICE described this three-state structure as a standard approach for estimating the cost effectiveness of cancer drugs and suitable for decision-making. The same structure recurs across tumour types and treatments.

Partitioned survival models have technical limitations. Their popularity also reflects available trial data, modelling convention, and computational convenience. They are not independent proof that a progression endpoint is valid in every setting. Still, the choice of states is revealing. When health economists must translate cancer into a tractable model, they do not divide the living population only by age or treatment status. They commonly divide it by whether the disease has progressed.

The same state logic appears in the Kesis Clinical applications that are in clinical testing. In the metastatic settings we model, governed medical knowledge developed through Nyx is organized around the disease state relevant to the current treatment decision. PFS helps describe the effect attributable to the treatment being considered, while OS, toxicity, quality of life, and other clinical consequences remain visible as separate considerations. The clinicians and other reviewers who have seen this design have consistently supported the approach. That reception is limited evidence: it shows familiarity with the state model, while application validity and outcomes remain separate questions. We have not had to teach people a new model of cancer for the structure to make sense.

Patients, clinicians, regulators, HTA committees, and application designers arrive at the same division for different purposes. The common intuition is that a cancer that remains controlled is meaningfully different from one that has escaped control, even when both people remain alive.

The boundary usually carries an implicit agreement

The 20 percent RECIST threshold is often criticized for being arbitrary. In one sense it is. There is no biological transformation at the moment the sum of measured diameters moves from 19 to 20 percent above its nadir. The underlying disease is continuous, heterogeneous, and only partly visible on scans.

But an operational boundary can be useful without being a law of nature. The more important question is whether the boundary matches the decision attached to it.

In a conventional trial of a new treatment for progressive metastatic cancer, people generally enter because the prior strategy is no longer adequate or because a new first-line treatment is needed. Randomisation assigns the new treatment strategy. The PFS clock then measures how long each assigned strategy keeps the cancer from reaching the next recognised loss of control. When progression occurs, treatment may change and the patient enters the much less standardised sequence of subsequent care.

This creates an implicit agreement. Before progression, the assigned treatment is ordinarily the treatment whose control is being measured. At progression, the disease state and treatment strategy may change together. The endpoint is useful because its state boundary and the clinical decision boundary broadly coincide.

Exceptions have always existed. Treatment can stop because of toxicity. A clinician may continue therapy beyond radiographic progression when the person is benefiting. Symptoms or organ risk can force a change before RECIST progression. Immunotherapy required modified criteria because apparent early progression can be misleading. These exceptions are handled through protocol definitions, sensitivity analyses, and clinical judgment. They do not erase the usual alignment on which PFS depends.

SERENA-6 tests what happens when the exception becomes the design.

A legitimate question that PFS may not answer

SERENA-6 asked an important and increasingly common question: should treatment change when molecular evidence of resistance appears, even though conventional imaging still says the cancer has not progressed? Earlier detection will force us to answer questions like this. The question is not the problem. The fit between the question and the endpoint may be.

The trial enrolled people with estrogen receptor-positive, HER2-negative advanced breast cancer receiving an aromatase inhibitor and a CDK4/6 inhibitor. Patients in whom circulating tumour DNA testing detected an ESR1 mutation, but who had no clinical or radiographic progression, could enter the randomised phase. One group switched the aromatase inhibitor to camizestrant and continued the same CDK4/6 inhibitor. The other continued the original combination. PFS was measured from randomisation to radiographic progression or death.

Among 315 randomised patients, median PFS was 16.0 months in the early-switch group and 9.2 months in the group that continued the aromatase inhibitor. The hazard ratio for progression or death was 0.44 (95 percent confidence interval 0.31 to 0.60). Those data answer a defined question: what happened to radiographic PFS when treatment changed at detection of an ESR1 mutation rather than continuing unchanged?

The clinically interesting question is slightly different. Is it better to use the new treatment at molecular detection, or to preserve the current treatment while the cancer remains controlled and use the new treatment at radiographic progression? SERENA-6 did not randomise that deferred strategy. Patients in the control group were not required to receive camizestrant with continued CDK4/6 inhibition when their cancer progressed. The FDA’s April 2026 briefing document identified the same missing comparison.

Consider two otherwise identical patients. Both have had meaningful control for a long time. Both develop the molecular signal. Neither has radiographic progression. Patient A changes treatment immediately and progresses 16 months later. Patient B remains on the current treatment for another nine months, then crosses the radiographic progression boundary, receives the same treatment Patient A received earlier, and remains controlled for another ten months. The numbers are illustrative, not SERENA-6 results.

The trial’s first PFS comparison would be 16 months for Patient A and nine months for Patient B. It would favour switching early. Yet Patient B would have had 19 months of control before exhausting the two treatment opportunities, compared with 16 months for Patient A. The conclusion reverses depending on where the clock stops.

Nothing needs to be falsified for this to happen. The scans can be read correctly, the randomisation can be intact, and the hazard ratio can be calculated perfectly. The design simply gives Patient A an active next-line treatment before the primary endpoint and stops counting Patient B at the moment that patient becomes eligible to receive it. If several available next-line treatments can delay progression, moving almost any one of them forward could create a first-PFS advantage. That would not establish that the selected treatment was uniquely effective, or that using it earlier was the better sequence.

This creates a way to game PFS without manipulating the endpoint analysis. Ordinarily, radiographic progression marks both the end of the measured state and the point at which the treatment strategy may change. In SERENA-6, the treatment decision moved to the earlier molecular boundary, while the primary endpoint remained attached to the later radiographic boundary. The trial treated the molecular state as important enough to justify a treatment change in one group, but still counted both groups as progression-free. One arm was allowed to spend a later treatment before the PFS event. The other arm’s clock stopped before its later treatment could count.

PFS2 does not automatically repair that asymmetry. In SERENA-6, the timing and choice of postprogression treatment were left to investigators, and crossover to camizestrant with continued CDK4/6 inhibition was not permitted. The FDA therefore described PFS2 as unreliable for determining whether treatment at molecular detection was better than treatment at radiographic progression. By the time PFS2 was measured, the two groups had received different numbers and types of new treatment without those later choices being randomised.

OS has the complementary problem described in that earlier essay. It remains a randomised outcome of the initial assignment, but everything that happens after progression enters the result. When later treatment selection is not controlled, OS cannot isolate whether the relevant difference came from using this treatment earlier, using a different sequence, unequal access to subsequent therapies, or the many patient factors that shape care after progression. PFS2 and OS may still describe what happened to the trial groups. Neither can retroactively create the sequencing comparison that was never randomised.

The clinical question called for a sequencing trial. A two-arm design could have compared switching at ESR1 detection with continuing the original treatment and then crossing over to the same new treatment at radiographic progression. The relevant later treatment could have been prespecified, and the endpoint could have extended across both sequences. If a no-crossover comparison was also required, a three-arm design could have included early switching, deferred switching to the same treatment, and a prespecified alternative sequence without crossover. That would separate the effect of treatment selection from the effect of treatment timing.

SERENA-6 is useful here not as a verdict on camizestrant or on early molecular switching. It is an example of a legitimate clinical question for which PFS may not be a useful endpoint. The analysis can be technically correct and the hazard ratio can be precise while the design rewards moving treatment earlier without showing that the resulting sequence is better.

The endpoint must get us closer to what patients want

Patients want to live longer and live better. Every clinical-trial endpoint is an attempt to learn whether a treatment helps them do one or both.

OS appears to answer the first question directly. As I argued in “Overall Survival Is the Weaker Endpoint”, however, survival after progression is shaped by subsequent therapies, access, crossover, comorbidity, disease biology, and clinical choices made long after randomisation. OS can be the correct outcome and still be a weak measure of what the treatment being studied actually did.

PFS can get us closer because the intuitive model of cancer control is connected to what patients want. We understand the direction without a statistical model: disappearance is generally better than shrinkage, shrinkage is better than no change, and no change is better than growth. Slow growth is generally better than rapid growth. Progression matters because growing cancer eventually causes symptoms, limits treatment options, impairs organ function, and threatens life. Keeping cancer on the controlled side of that boundary is therefore a plausible way of helping someone live longer or live better, even when the size of that effect cannot be predicted from PFS alone.

PFS gives that intuitive model a clock. It asks how long the cancer remained controlled before all the later treatments and patient factors entered the result. That is why PFS can answer something clinically important even when it does not reliably predict OS.

But PFS does not matter simply because it is PFS. The appraisal should be straightforward: given the way this trial was designed, does its PFS result get us closer to knowing whether patients live longer or live better?

In a conventional trial, delaying a meaningful loss of cancer control may do exactly that. In SERENA-6, first PFS did not resolve whether switching treatment earlier created more total control or merely moved the next treatment forward. The design could produce a longer first PFS even if the complete early-switch sequence was no better than waiting. For that clinical question, PFS was probably the wrong outcome selection.

The sequencing question required an outcome that extended across the relevant treatment sequence, with subsequent treatment prespecified, together with direct measurement of how patients felt and functioned. OS could still contribute, but it could not repair an uncontrolled sequence after the fact. The right outcome was whichever combination allowed the trial to distinguish additional time living well from the same treatment benefit being spent earlier.

The same rule applies beyond PFS. Response rate can be a valid endpoint when tumour shrinkage is the effect that matters. Time to symptom deterioration can be valid when preserving how a person feels is the question. PFS can be valid when delaying progression is a meaningful and fairly measured period of control. OS can be valid when the treatment strategy’s effect on survival can be interpreted. No endpoint is made valid by its position in a hierarchy. It is valid when it answers the question the trial needs to answer and moves us closer to knowing whether patients live longer or better.

That is our position on PFS. The intuitive model explains why delaying progression can get us closer to knowing whether patients live longer or better. Its imperfect surrogacy with OS does not make that connection meaningless. But the design of each trial must preserve it. When a PFS result reflects a genuine extension of cancer control, the endpoint can answer something patients care about. When the design can improve PFS without improving the patient’s overall period of control or experience, the connection has been broken and the endpoint was the wrong choice.

In clinic, when a patient asks whether treatment is working, the answer begins with the cancer: gone, smaller, stable, or growing. The immediate goal is intuitive. Eliminate it if possible, shrink it if we can, and otherwise keep it controlled for as long as possible. PFS gives that goal a clock. That is the endpoint we had already chosen.