← Back to blog

P10/P50/P90 Forecasts That Pinpoint Site Bottlenecks in Clinical Trials

October 3, 2026
P10/P50/P90 Forecasts That Pinpoint Site Bottlenecks in Clinical Trials

For interim monitoring and early planning, use Bayesian accrual models or parametric Poisson models; when you have large historical datasets, layer in machine learning or resampling methods for sharper precision. Start today by running a baseline forecast with P10/P50/P90 bands, then set a weekly or monthly cadence that feeds observed per-site actuals back into the model. Report uncertainty, not a single date, and recalibrate as real enrollment data comes in.


TL;DR:

  • Large historical datasets enable machine learning and resampling methods to refine enrollment predictions, but they require extensive site-level data for accuracy.
  • Accurate forecasting depends on input quality, including site activation dates, staffing capacity, seasonality, and clear categorization of screening failures, which should be regularly updated.
  • Combining study-level and site-level forecasts provides operational insight, with the former guiding overall timelines and the latter identifying sites needing immediate intervention.
  • Most forecasting errors stem from input assumptions and early data instability, making wide prediction intervals and regular recalibration essential for honest planning.
  • Operational constraints, such as IRB delays or staffing gaps, significantly impact enrollment timelines, and integrating diagnostics like HaiPhai helps address these bottlenecks effectively.

Haiphai
Turn Forecasts Into Faster Trials
HaiPhai helps life sciences teams identify operational bottlenecks and integrate AI into processes such as regulatory drafting and clinical site activation.
Explore HaiPhai

Table of Contents

Overview of forecasting approaches used in clinical trial enrollment

Every forecasting method trades off simplicity against data appetite. Bayesian accrual models combine a protocol's assumed enrollment rate with observed interim data, producing not just a predicted completion date but a full probability distribution around it. This is why the Accrual Prediction Program built at a cancer center relies on them for ongoing trial review. Poisson and other parametric processes are simpler and easier to explain to a steering committee. This makes them a reasonable starting point when a trial has little or no accrual history yet.

Machine learning methods, including gradient-boosted models and zero-inflated or hurdle models, need more fuel: site-level features and a large library of past trials. Given that, they can predict monthly site-level enrollment counts with notable gains over historical-rate baselines, as shown in machine learning research on enrollment prediction. Non-parametric resampling and Monte Carlo simulation sit at the flexible end, handling seasonal dips, pauses, or irregular enrollment curves that break the assumptions parametric models depend on.

Picking among them comes down to three questions:

  • How much historical, site-level data do you actually have on hand?
  • Do stakeholders need a transparent, auditable model or just a fast estimate?
  • Does the trial's enrollment pattern look steady, or irregular and disrupted?

Weighing accuracy, transparency, and data needs against each other

Choosing a method is less about which one is "best" and more about which trade-offs your trial and your oversight committee can tolerate. Site-level historical data, when available, makes machine learning methods viable and often more accurate; without it, parametric or Bayesian approaches built on study-level assumptions remain the more practical default.

Output shape matters as much as accuracy. A model that returns a single predicted date hides the risk; a model that returns a full predictive distribution, as Bayesian and simulation methods do, lets a project team make a risk-aware call about whether to add sites now or wait another month. The eventPred vignette illustrates this well: with sparse interim data, its time-decay and Weibull models produce genuinely wide prediction intervals rather than false precision, which is the honest answer when so little has happened yet.

Interpretability also shapes what gets approved. Regulators and data monitoring committees are often more comfortable with transparent parametric or Bayesian models than with black-box machine learning, unless that machine learning has been validated thoroughly against historical data.

A prospective validation study reported an AUC of 0.737 for a model predicting trial accrual failure from clinicaltrials.gov features. That is a meaningful signal for flagging at-risk trials early, though it is well short of a deterministic answer, which is exactly why prediction intervals matter more than point estimates.

When vetting any method, ask for:

  • MAE or MSE at the study level, benchmarked against a simple historical-rate baseline.
  • AUC or similar metrics for any binary failure-prediction task.
  • Calibration checks confirming that the model's P10/P50/P90 bands actually cover outcomes at the stated rate.

Operational inputs that change forecasts

A model is only as good as what feeds it, and enrollment forecasting runs on operational data that most CTMS systems already capture but rarely structure well.

  1. Site activation dates and their gating milestones: contracting, IRB approval, and kit or supply shipment. Model these as a distribution of likely start dates per site rather than a single fixed date, since activation timing is itself uncertain.
  2. Coordinator capacity in FTEs alongside screening funnel metrics: how many patients are screened, how many fail screening, and how many drop out after enrollment, captured per site rather than averaged across the study.
  3. Seasonality and competing trials: enrollment by calendar month, plus any local external factors such as competing studies drawing from the same patient pool.
  4. Standardized, machine-readable fields in the CTMS or feasibility log, so these inputs can feed periodic model updates without manual reformatting every cycle.

Pro Tip: Log screen-fail reasons as structured categories, not free text; that single change makes dropout and screen-fail rates usable as model inputs instead of a narrative nobody can quantify.

Practical forecasting workflow: plan, run, recalibrate, and act

A forecast that is never updated is a guess with extra decimal places. Build the workflow around a repeatable cadence from day one.

  1. Baseline planning: choose a method appropriate to your data depth, document every assumption in writing, and produce P10/P50/P90 estimates plus at least one sensitivity scenario.
  2. On-study cadence: update weekly for fast-enrolling trials or monthly for slower ones, feeding in per-site actuals and current activation status each cycle.
  3. Calibration: rescale priors and model parameters against observed actuals, and validate using both a retrospective holdout period and prospective tracking as new data arrives.
  4. Trigger rules: set explicit thresholds, for example a P50 projection missing the target enrollment date by more than 15% at a defined checkpoint, that automatically escalate to a mitigation conversation.
  5. Documentation: maintain a short audit trail of assumptions and changes for oversight committees and regulatory briefings.
Workflow stagePrimary questionOutput
Baseline planningWhich method fits the data and governance needs?P10/P50/P90 forecast plus sensitivity scenario
On-study updateWhat changed since the last cycle?Revised per-site and study-level projections
CalibrationDo the bands still match reality?Rescaled model parameters
Trigger and mitigationHas a threshold been breached?Decision to add sites or escalate resourcing

Tools and reference implementations worth knowing

Start with something reproducible before anything goes into production. R's accrual-focused packages, including the eventPred approach referenced earlier, let a biostatistician validate a method against historical data before trusting it on a live trial. Browser-based enrollment simulators serve a similar early-stage purpose: the open-source Trialsim simulator runs Monte Carlo projections with per-site ramp-up and P10/P50/P90 bands, and its documentation explicitly frames early forecasts as sketches meant to be recalibrated from actuals, not final answers.

Enterprise, CTMS-integrated forecasting becomes necessary once a forecast needs to support regulator-facing decisions or multi-study governance, since that context demands audit trails and version history that a standalone script cannot provide on its own.

  • Reproducible statistical packages for method validation before committing to a live model.
  • Browser-based simulators for fast, early feasibility trade-offs across site scenarios.
  • CTMS-integrated forecasting for ongoing, auditable, regulator-facing updates.
  • A staged path: validate with a script or package first, then operationalize inside the system of record.

How HaiPhai connects forecasting to site activation

Forecasts only matter if the operational levers behind them move. HaiPhai works as an embedded operational partner rather than a software vendor, starting from a company's strategic enrollment goals and working backward through its AI Velocity Diagnostic to find where activation actually stalls, whether that is contracting, IRB review, or staffing gaps. Rather than deploying a generic forecasting tool, this approach backtracks from the target timeline to pinpoint the specific bottleneck slowing a given trial, then builds the fix around that trial's own data and workflows. Pairing that diagnostic with a site readiness checklist gives project teams a concrete way to close the gap between a forecast's uncertainty bands and the activation delays actually causing them.

Why site-level and study-level forecasts need each other

A study-level forecast tells a sponsor when the trial as a whole is likely to finish enrolling; a site-level forecast tells a project manager which specific sites are falling behind and need intervention now. Relying on only one of the two leaves a real gap. A study-level model can show enrollment tracking on target in aggregate while masking the fact that three sites are badly behind and two are carrying the whole trial, a pattern that only shows up once you disaggregate.

Study-level and site-level forecast comparison

Site-level forecasts matter operationally because mitigation decisions, adding a site, reallocating a coordinator, escalating an IRB follow-up, happen at the site, not at the study. But site-level models also need more data to stay reliable: fewer patients per site means noisier estimates, which is exactly why methods like the Accrual Prediction Program pair site-level tracking with study-level priors to stabilize early estimates.

The practical answer is to run both and reconcile them regularly. When the sum of site-level projections drifts from the study-level forecast, that gap itself is informative: it usually means a handful of sites are driving results that an aggregate number hides. Tracking site performance against study-wide KPIs, including throughput and screen-fail rates, keeps both views honest and catches drift before it becomes a missed milestone.

Where enrollment forecasts commonly go wrong

Most forecasting failures are not modeling failures, they are input failures. Treating a protocol's assumed enrollment rate as fact rather than a prior is one of the most common mistakes: that rate is almost always optimistic, set during protocol design before any site has opened.

Sparse early data is another recurring trap. In the first weeks of a trial, almost any model produces unstable estimates, and teams who anchor on a single point forecast instead of a wide interval end up surprised when reality drifts. The eventPred vignette demonstrates this directly, showing genuinely wide prediction intervals when few events have been observed, which is the model being honest rather than being wrong.

Overreacting to short-term noise causes the opposite problem: a single slow month triggers a panic-driven mitigation plan when the underlying trend has not actually shifted. Dynamic linear models with weakly informative priors are designed specifically to smooth these estimates over time so teams respond to trends rather than noise.

Finally, forecasts built only from study-level aggregates miss the site-specific drivers, IRB delays, coordinator turnover, local competing trials, that actually explain most variance. A forecast without a feedback loop to recalibrate from actuals is the most common pitfall of all: it becomes stale the week it is built.

Feeding recruitment strategy into the forecasting model

A forecast is not just a passive prediction, it is a planning tool that should absorb every recruitment lever a team pulls. When a team adds a new site, launches a local advertising push, or expands eligibility criteria, that change belongs in the model as an updated input, not as a separate note sitting outside it.

Dropout and screen-fail rates deserve the same treatment. Separating screened, failed, and enrolled counts per site lets a forecasting model account for the recruitment funnel's actual shape rather than treating enrollment as a single clean rate.

Recruitment strategy changes should also trigger a forecast refresh on their own, independent of the regular cadence. Adding three sites mid-trial is a structural change to the model's inputs, not noise to be smoothed over at the next scheduled update. The practical habit is to treat every recruitment decision, a new site, a new patient-facing campaign, a protocol amendment loosening criteria, as a model input event, logged alongside the operational data described earlier so its effect on enrollment can actually be measured afterward instead of assumed.

What forecast-driven course corrections look like in practice

The value of a forecast shows up at the moment a team acts on it. A Bayesian accrual model that flags a declining probability of hitting a target enrollment date gives a project team a specific, quantified reason to open two additional sites rather than a vague sense that things feel slow. That is the entire premise behind the Accrual Prediction Program: it converts interim data into a predicted completion date, a prediction interval, and a probability of meeting the original target, giving oversight committees a concrete trigger rather than a gut call.

Weighted resampling methods offer a different kind of course-correction signal. Peer-reviewed validation on real recruitment logs found that weighted resampling approaches outperformed Bayesian baselines in scenarios where early enrollment patterns were irregular, achieving better coverage and smaller error against actual outcomes. A trial with a choppy early curve, a slow start followed by a ramp after additional sites activate, is exactly the case where this method's flexibility pays off over a smoother parametric assumption.

Dynamic linear models show their value in the opposite situation: steady trials where a single bad month should not trigger an overreaction. Smoothing quarterly accrual estimates with weakly informative priors, as documented in recent Bayesian modeling work, keeps project teams from escalating mitigation on noise while still catching a genuine downward trend early enough to act on it.

What forecast-driven course corrections look like in practice — overview diagram

Stop looking for a perfect model

Most project managers spend too long hunting for the one model that will nail the date. A defensible method with honest uncertainty beats a precise-looking model with hidden assumptions every time, because the uncertainty is the information that lets a committee make a real decision. Run a baseline forecast this week, show the P10/P50/P90 bands openly, and let operational fixes, not a better algorithm, do most of the work of closing the gap.

— John

Turning a forecast into fewer delays with HaiPhai

A forecast tells you a trial is behind. Fixing it is a different problem, and that's where most teams lose months they can't get back. HaiPhai works as an embedded operational partner rather than another dashboard: its AI Velocity Diagnostic maps where your specific trial is actually stalling, whether that's IRB turnaround, coordinator capacity, or contracting delays, and then redesigns the workflow around your team's own data rather than a generic template.

Haiphai

That matters because the gap between a forecast's P50 date and your approval timeline is rarely a modeling problem. It's usually a handful of operational bottlenecks that compound across sites.

  • An operating partnership can embed senior expertise directly into clinical and regulatory operations rather than relying solely on software.
  • The solutions built around institutional knowledge and governed automation extend that diagnostic into ongoing clinical, regulatory, and executive operations.

If forecasting has shown you where enrollment is at risk, talk with HaiPhai about mapping the operational fix.

Sources

These sources back the methods and figures covered above, drawn from peer-reviewed validation, institutional implementations, and open-source tooling rather than vendor marketing.

FAQ

How do you calculate the enrollment rate?

Enrollment rate is typically the number of patients enrolled divided by the number of active site-months, giving patients per site per month. Models like the Accrual Prediction Program use this rate alongside protocol assumptions to generate a predicted completion date and prediction interval rather than a single fixed figure.

What is a 3+3 trial design?

A 3+3 design is a dose-escalation method used in early-phase oncology trials, where small cohorts of three patients are enrolled at each dose level before deciding whether to escalate, de-escalate, or stop. It is a dose-finding framework rather than an enrollment forecasting method, though its small cohort sizes make enrollment timing especially sensitive to site-level delays.

Trial teams are increasingly pairing forecasting models with operational diagnostics rather than treating enrollment prediction as a standalone statistical exercise. Master protocols that share infrastructure across arms or trials are also drawing more attention for difficult recruitment settings, though the FDA's guidance on master protocols notes they add coordination complexity even as they improve efficiency.

How long does it typically take to go from phase 3 to market?

Timelines vary widely by therapeutic area, regulatory pathway, and how enrollment and trial execution unfold, so there is no single industry-wide figure to cite. Enrollment delays in phase 3, often driven by site activation and screening funnel issues rather than the statistical model itself, are frequently a major driver of how long that path actually takes.

How does patient dropout affect enrollment forecasts?

Dropout and screen-fail rates change the effective number of patients a trial needs to screen to hit its enrollment target, so a forecast that ignores them will understate the timeline. Tracking screened, failed, and enrolled counts separately per site, rather than a single blended enrollment number, lets a forecasting model account for the funnel's actual shape.