zahrcast
the sia model · interim decisions in oncology

bayesian forecasting for oncology trials in flight

Sia is a Bayesian hierarchical model. It joins a mechanistic model of tumour dynamics to a multistate survival model, and partially pools information across historical trials. Applied to a trial with short follow-up and most patients still censored, it forecasts the readout that trial will report at maturity, with credible intervals.

𓀁the problem

deciding on immature data

Interim decisions are made a few months into follow-up, while the cohort is still enrolling and most patients are censored. The mature survival curve the decision depends on does not exist yet.

The standard response is to wait for more events. That has a cost: enrolment continues during the wait, and in some cases the eventual answer was already implied by the tumour measurements on hand.

Waiting is paid for in trial time. A single pivotal trial supporting a new approval costs a median of about $19m to run,1 and the program it belongs to a median of about $985m per approved drug on public filings.2 Every month a decision waits on events is spent at that rate, and enrolment continues while it is spent.

So the question is how much of the mature answer is already implied by the measurements collected so far, and how confidently that can be stated.

what you get

Sia produces the full range of readouts this trial could still report, weighted by how plausible each is given the data collected so far. The go/no-go question is asked of that range directly: the probability that the benefit exceeds the threshold you care about, stated today rather than when the events have accumulated.

The same output answers the neighbouring questions — whether to continue or stop an arm, and how to size the next study — because they depend on the same unobserved readout.

Borrowing from a historical trial is an assumption, and how much is borrowed is estimated rather than fixed in advance: where the trials disagree, the pooling loosens and the intervals widen.

𓋷the model

tumour dynamics, then events

Most interim forecasting models the endpoint directly: events accumulate, a hazard is estimated, the curve is extrapolated. Sia models each patient's tumour burden trajectory instead, and derives progression and death from that trajectory.

This is why early data carries information. A patient with four recorded assessments contributes little to an event count but a good deal to an estimated growth curve.

1 · fit

The model is fitted to the data collected so far, together with the historical trial it partially pools from.

2 · roll forward

Each patient is simulated onward from where they actually are today: the assessments they have not had yet, and the progression or death those assessments imply.

3 · read off

Repeated thousands of times, this produces thousands of complete finished trials — every way this trial could plausibly turn out, given what has been seen so far. Every reported quantity — a KM curve, a median, a response rate, an arm difference — is read off that set. (Formally: draws from the posterior predictive distribution.)

So “a 70% probability the benefit exceeds two months” means it in the plainest sense: in 70% of the simulated trials, it does. Nothing here is a single prediction with error bars attached. The answer is the whole set of outcomes, and it is wide for two reasons: patients differ, and the data so far leaves room for more than one version of the model.

patient-level inputs forecast endpoints tumour measurements sum of longest diameters baseline covariates clinical, biomarker, ctDNA event history progression, death, dropout mechanistic submodel tumour burden over time: decline, nadir and regrowth, per patient latent time-varying covariates current burden and growth rates multistate submodel progression-free progressed off-trial deceased progression-free survival composite and cause-specific overall survival routed through the states other endpoints response rate, hazard ratios, landmark survival

The two submodels are fitted together, not in sequence: the tumour trajectory supplies the covariates the event model runs on, and the event data in turn constrains the trajectory. Progression, death and off-trial follow-up are states in one structure, so progression-free and overall survival come from the same simulated trials and cannot contradict each other — and dropout enters the forecast rather than being assumed away. Redrawn for this page from the manuscript's own diagrams, with the notation left out.

fig_sld_trajectoriesassets/figures/fig_sld_trajectories.svgraw tumour data — decline, then regrowth

Sum of longest diameters over time, one line per patient. Decline followed by regrowth is the common shape. Records stop early and at different times, which is what the censoring looks like before modelling.

fig_sld_censoredassets/figures/fig_sld_censored.svgforecast continuation of an in-flight trial

The same patients continued forward. Observed measurements are on the left of each panel and the forecast of the unmeasured remainder on the right. Forecast width scales with how little was observed: a patient with two visits gets a wide interval.

𓏛case study

two first-line es-sclc trials, n = 497

The case study pools a historical first-line extensive-stage small-cell lung cancer trial with a target trial, then asks the target trial's question at three cut-offs during its conduct. At each cut-off the model is refit from scratch using only data available at that date. No later data enters the fit.

fig_trial_timelineassets/figures/fig_trial_timeline.svgswimmer plot with the three cut-off markers

One bar per patient, from enrolment to last recorded assessment, ordered by enrolment date, so the staircase shape is the accrual curve. Ticks mark recorded visits. The dashed lines are the three forecast cut-offs.

cut-offmonthtarget patients observedwhat it is
first49earliest forecast; interval already covers the mature curve
second1139the month the OS forecast settled on the mature answer
third1971mature reference — the month the trial's own data agreed

Thirty-two patients enrolled between the second and third cut-offs. The eight-month gap between those cut-offs is what happened on this trial, not a rate to expect elsewhere. It is also a lower bound on the lead time: the trial's actual decision came later, after database lock and full follow-up.

provenance

The case study uses de-identified patient-level data from two first-line ES-SCLC trials, a Lilly CXCR4 trial and Amgen 20010145, obtained through Project Data Sphere. The trials are named here because the manuscript names them. The Sia model core is licensed under the PolyForm Noncommercial License 1.0.0: free for research and other noncommercial use, with commercial use licensed separately by zahrcast.

𓆄convergence & calibration

forecasts checked against the reported result

Each column below is a separate fit, using data up to its cut-off only, forecasting the curve the trial eventually reported. Two things are worth checking across the columns: whether the credible interval covered the mature curve at every cut-off, and whether it narrowed as data accumulated.

fig_lfo_pfsassets/figures/fig_lfo_pfs.svgforecast evolution across cut-offs — the central result

Progression-free survival at the three cut-offs and on full data. At month 4, with nine patients observed, the credible interval covers the mature month-19 curve. The interval is wide at that cut-off, which is the expected behaviour with nine patients.

fig_lfo_osassets/figures/fig_lfo_os.svgoverall survival, same cut-offs

Overall survival is the harder endpoint and the intervals are wider throughout. The forecast converges on the mature result by month 11, with 39 patients observed. Deaths are sparse in an immature trial, and this is the limit that sparsity imposes.

Three limits. Early intervals are wide. No single cut-off here would have justified a go/no-go decision on its own. And a result on one disease and two trials does not transfer to another without checking. Converting a forecast into a decision requires a benefit threshold and a cost of being wrong, both of which are yours to set.

𓂋contact

if you have a trial in flight

A first conversation is best about a specific trial: what is enrolled, what is being measured, when the decision is due, and what would change if the answer arrived earlier. No data sharing is needed for that conversation.

reading

Manuscript — the ES-SCLC forecasting study, with the full leave-future-out results and calibration checks. arxiv.org/abs/2607.17908.

Sia — the model core: hazards, GP knot grids, tumour dynamics, the multistate likelihood and leave-future-out scoring. The source is private, licensed under PolyForm Noncommercial 1.0.0. Access and commercial terms by arrangement.

references
  1. Moore TJ, Zhang H, Anderson G, Alexander GC. Estimated costs of pivotal trials for novel therapeutic agents approved by the US Food and Drug Administration, 2015–2016. JAMA Internal Medicine. 2018;178(11):1451–1457. doi:10.1001/jamainternmed.2018.3931.
  2. Wouters OJ, McKee M, Luyten J. Estimated research and development investment needed to bring a new medicine to market, 2009–2018. JAMA. 2020;323(9):844–853. doi:10.1001/jama.2020.1166. Median estimate $985.3m, reported with a 95% CI of $683.6m–$1,228.9m. (See the 2022 correction notice, JAMA. 2022;328(11).)