Insights · FP&A and forecasting

The forecast you can argue with

Patient-based modelling, and what it is really for

A patient-based model is uncomfortable to build, because it forces every assumption it rests on into the open where somebody can argue with it. That discomfort is the whole point.

I am writing this from a pharmaceutical and life-sciences perspective, because that is where the method was developed and where the cascade is most explicit. The skeleton travels a long way beyond it, and I come back to that at the end, but the detail below is pharma detail and it is meant to be.

Where trend lines fail

A trend line can become an answer with the reasoning deleted. It takes the shape of the past and pushes it forward, and it flatters the person presenting it, because the arguments it invites are about method rather than about the business. You can contest the period, the seasonality treatment or the handling of outliers, and a good analyst will. What you cannot do is point at the commercial assumption carrying the number, because the method never put one on the page. So when the forecast misses, and in any business with real dynamics it eventually does, there is nothing commercial to take apart.

A patient-based model works the other way round. It refuses to start from the revenue line, and builds the number instead from the thing that causes it. Every step from the population down to the pound is a separate figure that somebody owns and can defend or contest. It is harder to build and much easier to attack. Both of those are features. The value of a forecast is not that it is smooth. It is that when it moves against plan, you can say exactly which assumption moved, and why.

The industry has a long-standing accuracy problem here, though it is worth being precise about what it does and does not prove. McKinsey's analysis of 210 launches between 2003 and 2009 found roughly two thirds missed pre-launch consensus analyst expectations in their first year. More recent work by Trinity Life Sciences puts it nearer half for the 2020 to 2023 US launch class, with a meaningful minority overperforming rather than under.

The honest reading matters, because the lazy one argues against itself. Those launches were not short of patient-based models. The method was already standard. What the misses tell you is that launch outcomes turn on access timing, uptake and competitive response, and that under-forecasting is as expensive an error as over-forecasting. Which is the argument for the method, not against it. A structured model is the only thing that tells you which of those assumptions broke. A trend line just tells you the pattern did.

Baseline, adjustment, and a word people get wrong

Before the mechanics, the vocabulary, because this is where the conversation usually derails.

The extrapolated number has a proper name: a baseline forecast. Data-driven, projected from history and available signals. It is not lazy and it is not wrong. It is incomplete, because what moves a pharmaceutical revenue line hardest is what has not happened yet. The work is turning that baseline into a judgementally adjusted, decision-ready forecast by layering on what the organisation knows and the data cannot hold: operational knowledge, commercial insight, planned initiatives, one-off events.

You will hear people say a model produces a "biased" forecast and human expertise makes it "unbiased". That is a misuse of a precise term, and it is worth being exact. In forecasting, bias is a systematic tendency to run high or low across many periods, and it is measured after the fact against actuals. You cannot know whether any forecast is biased until you have watched it for several cycles. A judgemental adjustment might improve accuracy or damage it, reduce bias or introduce it. You find out afterwards.

Three concepts, related but distinct, and routinely collapsed into one:

  • Accuracy. How close the forecast landed to reality.
  • Bias. Whether it consistently over- or under-predicts.
  • Judgemental adjustment. Human intervention using knowledge the model does not hold.

So the objective is not removing bias. It is improving forecast quality and decision usefulness, then measuring accuracy and bias against actuals over time to find out whether the judgement earned its place. That test has a name in the profession, forecast value added: did the adjustment beat the baseline it replaced?

The anatomy of the model

The pharmaceutical version is the clearest teaching example, because the cascade is explicit. The stage names below are the convention I use. They are not an industry standard, and different houses split eligible, addressable, accessible and target population differently. That is worth stating rather than glossing, because half the arguments about a patient model turn out to be arguments about where somebody drew a line.

You start with epidemiology: how large is the patient population, prevalent or newly incident each year, split where it matters by stage, subtype or biomarker. From there you step down through the diagnosis rate, because only diagnosed patients enter the funnel, and underdiagnosis is often the largest single number in the model. Then the treatment rate, because not every diagnosed patient is treated.

Then comes the stage most summaries skip, and it is the one that decides whether a model is credible. The eligible population is not everyone treated. It is those who qualify for this specific product under its label: the right line of therapy, the right biomarker, the right severity, the right prior therapy, no disqualifying contraindication. And eligible is still not addressable, because a patient who qualifies clinically but sits behind a reimbursement restriction or a step edit cannot be reached. Addressable, after label and after access, is the market. Not the diagnosed population, and not the disease.

PATIENT-BASED FORECASTING Revenue is the answer at the bottom. Not the input at the top. Population who is out there Prevalence / incidence who has the disease Diagnosed underdiagnosis is oftenthe biggest number Treated on any therapy Eligible qualifies under YOUR label Addressable reachable after access.THIS is the market Share of the DYNAMIC pool,not the whole Volume persistence x intensityx dose Net revenue gross-to-net, never list Why compliance varies, and why one rate across a portfolio fails They forget. Age, cognitive decline. Capacity, not choice. Persuasion does nothing. They feel sick. Punishing therapy in advanced disease. Shows up as persistence. They feel better. They stop. The relapse arrives months later, invisible. A board does not buy a forecast. It buys the ability to interrogate one. jatinderpurewal.com
The patient cascade: revenue is the answer at the bottom, not the input at the top

Only then does share apply, and it applies to the dynamic pool rather than the whole of it: newly diagnosed patients, and those switching from a prior therapy. The existing stock of patients stable on a competitor is sticky, and a model that applies a share percentage to the total addressable population has quietly reverted to being static. Share itself is shaped by the uptake curve, order of entry, competitor launches and, eventually, erosion at loss of exclusivity.

Converting patients into units is where the terminology gets muddled, including in places that should know better. The current framework separates three things:

  • Initiation. Whether a prescribed patient ever starts. Primary non-adherence and abandonment at the pharmacy counter are large, measurable and commercially actionable, particularly for high-cost specialty products, and they are the element most often left out entirely.
  • Implementation, usually called adherence or compliance. How completely a patient who is on therapy takes the prescribed dose. This drives intensity, not duration.
  • Discontinuation, measured as persistence. How long a patient stays on therapy before stopping. This is a survival curve, and persistence is duration of therapy, not an input to it.

That last distinction is where most write-ups go wrong, including an earlier version of this one. A patient can persist twelve months at seventy per cent dose fidelity: the duration is still twelve months, the consumed volume is not. So volume runs as persistence multiplied by dose intensity multiplied by dose per day, and the two variables have entirely different commercial fixes.

And underneath all three sits the regimen itself, which is a modelling input rather than a clinical footnote. What is actually prescribed, how often, by what route, in what cycle, with what titration and what treatment holidays, is what converts a patient-month into packs. It is also the single largest determinant of whether the patient complies at all, and the size of that effect surprises people who have not modelled it.

The published evidence is unambiguous, and the figures have to be compared like for like or they mean nothing. In a meta-regression of chronic cardiovascular medications, regimen adherence ran at 83.1% once-daily against 73.8% twice-daily; timing adherence, the stricter measure, ran at 74.2% against 50.4%. After adjustment, twice-daily dosing was worse by 14.2 points on regimen adherence and 22.9 points on timing. These are pooled across studies and populations, not a within-patient crossover, but the direction and the magnitude are consistent across the literature. Twenty-three points of timing adherence is not a modelling nicety.

And the regimen is only half of it, because the reasons a patient does not comply are not one thing. The literature splits them into unintentional and intentional, and the distinction is not academic: they have completely different commercial answers.

Unintentional is capacity. An elderly patient, or one with cognitive decline, forgets. So does anyone on four other medicines. Nothing here is a decision, so nothing here is fixed by persuasion. It is fixed by simplifying the regimen, by packaging, by reminders, by involving a carer.

Intentional is a judgement the patient has made, and it runs in both directions. In advanced disease, particularly metastatic cancer, the therapy can be punishing. A patient who feels sick will miss doses, and past a certain point will stop altogether, which shows up in your model as persistence rather than implementation and needs a tolerability answer rather than a support one. The mirror image is just as common and far more easily missed: the patient feels better, decides they no longer need it, and stops. The relapse arrives months later, by which point the discontinuation is invisible in the data that caused it. Anyone forecasting an asymptomatic chronic condition, where the medicine produces no felt benefit at all, is forecasting against exactly this.

The framework that captures it is Horne's necessity-concerns balance: adherence tracks how strongly a patient believes they need the medicine against how worried they are about taking it. Feeling better lowers perceived necessity. Side effects raise concerns. Both land in the same place in your volume line, and neither is visible in a dosing table.

This is why an assumed compliance rate is one of the most dangerous single cells in the model, and why one rate across a portfolio is indefensible. The number is a function of the disease, the symptom profile, the tolerability and the patient population, not of the molecule. Plug in ninety per cent because it sounds prudent, when the real-world figure for a twice-daily oral in an asymptomatic condition is closer to sixty-five, and you have overstated volume on that product by a quarter, inside a waterfall that looks rigorous all the way down.

And because the regimen is frequently a design choice rather than a given, this is one of the places finance should be in the room early, at target product profile stage, rather than receiving the assumption later and being asked to forecast it.

The whole set is also indication-specific. In oncology, time on treatment is driven by progression and measured directly, adherence applies mainly to orals, and dosing is frequently by weight or body surface area rather than a flat strength. A single "compliance" cell can be perfectly appropriate in a small commercial model, provided the parameter is defined and is not already baked into the observed utilisation you calibrated against. Double-counting non-adherence that the dispensing data already contains is a more common error than omitting it.

Price is not a bottom-line adjustment, and it is never list. The gross-to-net bridge carries rebates, chargebacks, confidential discounts, tender outcomes, patient support and statutory clawbacks, and in many markets it is where a large and growing share of the value now sits. Two live examples make the point better than the principle does. In the UK, the VPAG payment rate on newer branded medicines fell from 22.9 per cent in 2025 to 14.5 per cent in 2026: an 8.4 percentage-point movement in the gross-to-net bridge on eligible UK NHS sales, driven by market-level factors including growth and patent expiries shifting products out of the newer category into the older one. Nothing an individual company did. In the US, the Inflation Reduction Act has put a negotiated maximum fair price and a redesigned Part D into every long-range model, with the nine-year small molecule and thirteen-year biologic asymmetry now shaping which assets are worth developing at all.

And price is not simply the last step. Access is bought with price concessions, so the addressable filter and the net price are jointly determined, not independent layers. Modelling them separately is a standard and expensive error.

And this is a multi-country problem, not a UK one. Germany has more than doubled its statutory manufacturer rebate. The GKV-BStabG came into force on 30 July 2026, adding an 8.5% markdown on top of the existing 7% under §130a SGB V and taking the total compulsory markdown on patent-protected medicines to 15.5%. That sits alongside Festbeträge, the reference-price system capping reimbursement for clusters of comparable products, and the AMNOG process, where free pricing (halved from twelve months to six in 2023) is followed by a negotiated price with the rebate backdated. Every developed market has some version of this machinery, and the versions do not agree with one another.

For anyone forecasting across a multi-entity group, that is the point: your gross-to-net bridge is not one bridge. It is a different one per country, moving on a different political clock, and a change in one market can propagate into others through reference pricing. My judgement, flagged as judgement: this only intensifies, and Germany at the end of July is the evidence rather than the exception. Governments are carrying deficits, populations are ageing, and healthcare costs are rising faster than the revenue base funding them. Mechanisms that hold down the drug bill get stronger, not weaker, and they propagate: one market finds a lever that works and others copy the structure. A statutory markdown, a reference-price cluster, a clawback formula, a negotiated maximum price. Different instruments, one direction of travel. Whether the current model survives that pressure unchanged is a bigger argument than this piece, and worth having separately.

Two further layers separate a real model from a diagram. Stock and flow: new starts and continuing patients behave differently, and an annual static funnel hides that. Monthly cohorts, and in oncology a multi-state structure moving patients between lines of therapy, capture the timing of revenue rather than only its size. And calibration, which is the step that separates a model from a diagram. You back-calculate implied treated patients from observed sales and reconcile that against what your epidemiology says. Doing it properly means a bridge, not a ratio: syndicated sales are ex-factory or ex-wholesaler and therefore include channel inventory, while prescription data is demand, so stocking, destocking and launch pipe-fill routinely explain a gap that looks like a modelling error. Free goods, patient assistance and compassionate use consume patients without generating revenue. Combination and concomitant regimens mean a class-level unit count is not a unique-patient count, so share across a class can legitimately exceed one hundred per cent.

When the two views disagree, it is usually not that one is wrong. It is that they are measuring different universes, periods or channels. The work is to explain the bridge and quantify what is left over, and it is the most useful afternoon in the whole build.

What it was actually for

I first built one of these in 2003, at Chiron Biopharmaceuticals, standing up long-range planning across Europe and the international markets alongside Corporate Strategy and Business Development, and replacing a plan that had been built by extrapolation. The method was well established in the industry by then. What was new was applying it to that plan. I built and maintained the model, pulled the sales and historical data, and worked with each function to fill in its piece of the assumption set, so the revenue line carried a structure underneath it instead of a single asserted figure.

Its value was in prioritising R&D and commercial investment against each product's real revenue drivers. Pharmaceutical R&D runs a decade or more ahead of revenue, and the products that reach market have to pay for the ones that did not. A patient-based model is what let us prioritise that investment with some rigour, rather than by instinct or by whichever therapy area argued loudest for its budget.

The model also gave direction, and this is the part that changed how the commercial teams worked. A forecast built from epidemiology down shows not only that a product missed its number, but which stage of the cascade moved. If treated patients were on plan but the product's share of that treated population had slipped, that points at market penetration, and it sends you to look at awareness and advocacy within the disease area before you spend anything. If share held but patients were starting and then stopping, that is a persistence problem, and patient support is a more likely answer than marketing spend. If they stayed on therapy but were not taking the full dose, that is adherence, and it is a different intervention again.

That diagnostic step is what fed sales, marketing and medical affairs with where to point their budget, product by product and country by country. We ran it at country level for individual products, including scenario work on a competitive threat to an established product and on different launch aggressiveness for a new one. It is worth being exact about what this method does and does not buy you, because the two get conflated constantly, including by me in an earlier draft of this piece. A patient-based model is a long-range instrument. Its job is to make the multi-year revenue line realistic, and the reason that matters is not forecasting elegance: resourcing follows the revenue line. If the long-range number is wildly off, headcount, manufacturing, launch investment and trial spend are all sized against a fiction, and the overspend is committed years before anyone notices. Getting that number as close to right as the evidence allows is what lets you resource a programme properly without overspending on it.

What it does not do is fix your short-cycle forecast accuracy or your close. That is a different discipline, and in the same role it was a different piece of work: automated bridge reporting built at profit-centre and cost-centre level. Top line driven by demand units and average selling price per unit. Cost of goods as a standard variable driven by unit across Europe. Cost centres by functional line, with medical affairs and sales and marketing also carried at profit-centre level, and general and administrative costs across the board. The engine was the template set: actuals upload, and the bridge against last forecast and against budget falls out automatically, so the business argues about drivers within days instead of rebuilding a spreadsheet for a fortnight. That is what took annual forecast variance to actuals from roughly fifteen per cent to roughly four, and the monthly close from five and a half days to three and a half.

Most of what business intelligence and AI tooling sells today is that same closed loop. It was templates then. The principle has not moved: the value is not the report, it is that the variance explains itself before the meeting rather than during it.

None of this guarantees accuracy. A patient-based model is only as trustworthy as the calibration behind each stage, and a seven-stage waterfall built on seven unexamined guesses is worse than a trend line, not better, because it now looks rigorous while it is still wrong. Over-layering is a real failure mode and it biases forecasts downward: every additional filter multiplied through the cascade shrinks the opportunity, and nobody notices, because each individual filter looked reasonable. The discipline only earns its keep if every stage is checked back against what happened and revised when it misses, and the depth of the model should match the weight of the decision behind it.

The label moves, and the market moves with it

Here is a dynamic that a static cascade misses entirely, and it is one of the most important things a long-range model has to carry.

Most products do not launch into their full market. They launch into a narrow, severe indication first, then extend the label outward as evidence accumulates. In oncology it is the common pattern: drugs are typically approved first in later lines of therapy, and nearly all products that go on to gain subsequent indications move from later lines toward earlier ones, and from advanced disease toward adjuvant and early-stage settings.

There are good reasons it works that way. In advanced disease the alternatives are exhausted, so the risk-benefit calculation tolerates a new agent with an incomplete safety picture. Trials read out faster because events accrue faster. And the comparator is weak, which makes a meaningful incremental benefit easier to demonstrate, which in turn is what carries the price through a health-technology assessment. Move earlier in the disease and every one of those gets harder: the standard of care is better, the bar for incremental benefit rises, and the trial takes longer because patients live longer.

The commercial consequence is that the addressable population is a step function, not a line. Each label extension opens a new and usually much larger pool, at a date you do not control, with its own uptake curve and its own access negotiation. Early-stage populations dwarf late-line ones.

Which means a competent long-range model is not one cascade. It is a stack of cascades, one per indication, each with its own eligible population, its own launch date and its own probability of ever happening at all. Model it as a single growing market and you will be wrong in both directions: too optimistic before the extension lands, and far too pessimistic after.

The valuation payoff

The discipline is at its most valuable where there is no history to extrapolate at all. When you value an in-licensing deal, an out-licensing arrangement or an acquisition, the number that decides it is a forecast of a product that may still be years from market, or may never reach it. A trend line cannot value something that has not launched. A patient-based model can, because the cascade can be built from epidemiology and analogues without a single unit of sales history.

The valuation that sits on top of it is a risk-adjusted net present value, and the common mistake is to build a full NPV and multiply the answer by one probability. That is not what rNPV does. Each cash flow is weighted by the probability that it happens at all: current-phase costs are close to certain, later development costs are conditional on the earlier phase succeeding, and launch revenues carry the full cumulative probability of technical and regulatory success through every remaining transition. The earlier the asset, the wider the gap between the two approaches.

Two things travel with that. First, PTRS is not probability of commercial success. Approval does not buy access, and the two failure modes need separating. Second, the discount rate has to be consistent with it. If development risk already sits in the probabilities, loading a venture-style discount rate on top counts the biggest risk twice. A deterministic approval-case NPV is not wrong in itself, provided nobody presents it as the expected value.

At GSK the underlying models were built by R&D and Business Development, and the commercial decision on whether to license a product out had already been taken by the business. I tested the commercial documentation, reconciled it, and built the case that went to the in-licensing and investment boards.

Where I actually earned my place was the second gate. A group can want a licence for perfectly good commercial reasons: category leadership in a market, a portfolio gap, a competitive block. But the legal entity that will hold the licence owes duties to its own company, and its priorities are not automatically the group's. That board has to be satisfied the licence is priced on an arm's length basis. Justifying that, using the commercial understanding of the product and the market rather than a transfer-pricing formula alone, is the part I was there for.

It is a subtle distinction and worth being precise about it: I was not the person deciding the deal, and I was not the person building the model. I was the person who had to make the economics defensible to a board with different interests from the one that wanted it.

What made it work was that every assumption in the underlying model was visible and separately negotiable, so a sceptical board could pull on any single thread: raise the diagnosis rate, cut the peak share, delay the launch by two quarters, and watch the valuation move in a way they could follow. By the time they approved it, they had stress-tested it themselves. That is a different thing from being persuaded.

It started in 2003. It did not stop there

The patient-based work began at Chiron in 2003. The discipline underneath it has run through everything since, applied to whatever unit actually caused the revenue. At GSK it was the assumption set behind licensing valuations, where the cascade had to survive a board that was not the one that wanted the deal. At Parexel it was site utilisation and contract costing: same logic, different unit, and the difference between recovering fixed overhead and not. At Shionogi and Accord it was multi-entity forecasting across markets that each behaved differently. And for the last eleven years it has been my own P&L, where the demand unit is an order, the forecast decides an inventory buy with my own capital behind it, and the feedback loop is immediate and unforgiving.

Different units, different data, the same argument every time: build the number from the thing that causes it, put every assumption where somebody can attack it, then check it against what happened.

What has changed since 2003

The skeleton has not changed: population down to net revenue, in stages. Nearly everything else has, including things that are genuinely part of the method rather than around it.

The data changed first. In 2003 you worked from syndicated epidemiology reports, prescription audits and physician surveys, and physician recall was doing more load-bearing work than anyone admitted. Longitudinal claims and electronic health records now let you observe the actual treatment pathway rather than infer it: who started, who switched, who stopped, and when. Persistence stopped being an assumption and became something you can measure.

The structure loosened. Static annual funnels have largely given way to monthly cohorts and, in oncology, multi-state models moving patients between lines of therapy. Analogue selection used to mean a few hand-picked comparators. It now means matching against large libraries of real product histories.

The output changed shape. A point estimate with a sensitivity table has become a distribution. Monte Carlo across the joint uncertainty of share, timing, access and persistence gives a board a range with the drivers named, which is a more honest object than a single number carried to three decimal places.

And the old binary went away. Patient-based versus demand-based used to be a choice. Best practice now runs both and reconciles them, using observed demand to calibrate the patient model and treating the residual gap as a diagnostic rather than an embarrassment.

What has not changed: Excel is still the lingua franca in most brand teams, and for good reason, because an auditable model somebody can open beats an elegant one they cannot. Rare disease and European patient-level data are still sparse. And human judgement on the target product profile and the competitive response still decides the story.

What AI changes is narrower than the noise around it suggests, and more interesting. The genuine capability is estimating a biomarker-eligible population from claims and records where the full population was never tested for the biomarker. That is a real gap in rare disease and oncology which no amount of analyst time could close, because the data was never collected. Alongside it, language models are now good at extracting structured epidemiology from unstructured literature, trial registries and HTA dossiers, which used to be weeks of manual reading.

The second thing it changes is cost. A living patient-based model used to be analyst-weeks, so it was rationed to the decisions that could justify the spend and most planning cycles never got one. That constraint has largely gone, which matters more for a mid-sized company than for a large one.

The limits deserve the same precision. These models are hard to explain to a board that has to sign the number, they drift, and European data protection and data sparsity bite hardest in exactly the rare-disease settings where the help would be most valuable. Most companies are still running this at pilot scale. My own working rule in a finance function is unglamorous and non-negotiable: model logic and aggregate inputs only, no patient-level or commercially confidential data leaving the controlled environment, version control and independent recalculation on anything that moves a decision, and a named human owner for every assumption. AI that cannot survive a change-control conversation does not belong in a forecast the board is going to approve.

Cheaper rebuilds also do not make a model more current, and pharmaceutical data shows why. Data vintage does not lag uniformly, and that is the part people get wrong. Retail prescription data is published weekly, though it is a projection from a sample rather than a census. Hospital channel data exists too and has for decades, but it typically arrives on a slower cycle and with different coverage by country, which matters for the hospital-administered oncology products most likely to need this treatment. The trap is not that hospital data is unavailable. It is assuming it refreshes on the same clock as retail. Claims split two ways: open, pre-adjudication data arrives in thirty to sixty days, while the closed adjudicated data you actually need for persistence averages around ninety days with a longer tail. Persistence is slower again for a structural reason, because you cannot measure a twelve-month window in less than twelve months. Epidemiology is slowest, with full cancer registration in England running eighteen to twenty-four months behind and the United States SEER public release landing roughly twenty-eight months after the year it describes.

The field is working on it, and a current model should use what exists: England's Rapid Cancer Registration Dataset was built precisely to bypass that wait, SACT gives treatment and line-of-therapy data, and SEER is implementing real-time reporting.

So a model refreshed weekly is current in its demand signal and stale in its denominators, and the staleness increases the further up the funnel you climb. The discipline that follows is to date-stamp every stage of the waterfall with the vintage of the data behind it, so nobody reads a two-year-old denominator as a description of today's market.

What this has to do with the price of medicine

People ask why medicines cost what they do. The honest answer starts with the failure rate.

On Citeline's analysis of programmes over 2014 to 2023, the phase-by-phase transitions run: about 47% from Phase I to Phase II, 28% from Phase II to Phase III, 55% from Phase III to submission, and 92% from submission to approval. Multiply those and you get the cumulative likelihood of approval from Phase I: 6.7%.

WHY MEDICINES COST WHAT THEY DO Of drugs entering Phase I, about 6.7% reach approval. Phase I to Phase II 47% Phase II to Phase III 28% where medicine goes to die Phase III to submission 55% Submission to approval 92% 0.47 x 0.28 x 0.55 x 0.92 = 6.7%. The chain reconciles to the headline. The trap, and it is a common one in risk-adjusted valuation Transition rate is the chance of reaching the NEXT phase. The bars above. Cumulative likelihood of approval is the chance of reaching MARKET from here. Chains and cumulative rates must come from the SAME study and period, or they will not reconcile. The optimistic story is access, not price Development efficiency does not lower what a medicine costs. Prices are not cost-recovery. What it changes is which medicines get made at all: smaller populations and worse-served indications become fundable when the expected cost of reaching approval falls. Source: Citeline / Norstella, programmes 2014-2023. Other datasets give different chains; each reconciles to its own cumulative. jatinderpurewal.com
Clinical development attrition: 47 / 28 / 55 / 92 multiplies to a 6.7% cumulative likelihood of approval

Note what that means. Phase II is where medicine goes to die, and that holds on every dataset I have seen. Every approved product carries the cost of all the ones that did not make it.

Two numbers get used interchangeably and are not the same: the transition rate between two phases, and the cumulative likelihood of approval from where you stand today. Confusing them is one of the most common errors in a risk-adjusted valuation, and it is not a rounding difference. Chains and cumulative rates must also come from the same study and the same period, or they will not reconcile. Mix two datasets and your own numbers contradict each other.

PTRS, which you will see throughout pipeline valuation work, is simply the probability of technical and regulatory success: the chance an asset clears the remaining development and regulatory hurdles. It is not the probability of commercial success. Approval does not buy access, and the two failure modes have to be modelled separately.

Which is why the discipline matters well beyond the finance function. A patient-based model is how a company decides which programmes to fund and how heavily to resource them. Resourcing follows the revenue line. If the long-range number is wrong, headcount, manufacturing and launch investment are all sized against a fiction, and the overspend is committed years before anyone notices it.

Now, the tempting conclusion is that development efficiency eventually brings the price of medicines down. It does not, and it is worth being straight about why, because it is the industry's least credible talking point. Drug prices are not cost-recovery. By the time a price is set the R&D spend is sunk, and the price is decided by value, unmet need, the competitive set and what payers will bear. That is a structural argument, not an empirical one, and it is why the US Congressional Budget Office states plainly that sunk R&D cost does not determine a drug's price. The empirical work points the same way: a study of sixty FDA approvals found no association between estimated R&D investment and treatment cost, at list price on launch or net price a year later. Sixty is a small sample and I would not hang the argument on it alone, but it is consistent with the structure and nothing credible points the other way. Saved development capital goes to the P&L, the pipeline or the shareholders. There is no pipe connecting it to price.

The real optimistic story is better, and it is about access rather than price. Development efficiency lowers the revenue threshold a programme has to clear to be worth funding. That does not change what an existing medicine costs. It changes which medicines get made at all. Smaller populations, worse-served indications, the rare diseases where the addressable pool was never going to support a blockbuster case: those become fundable when the expected cost of reaching approval falls.

That is something a finance leader genuinely influences, and it is worth more than a slogan about prices.

Where else this works, and where it does not

The skeleton is not pharmaceutical. It is what you build whenever revenue depends on a population that has to be acquired, converted and then kept. The closest neighbour is any consumable or repeat-purchase business, and the mapping is tighter than it looks: aware and in-market stands in for diagnosed, fits-the-product-and-price-point for eligible under label, reachable-through-channels-you-can-buy for addressable after access, share of new and switching buyers rather than the installed base, and then the first repeat, which is the direct analogue of initiation and almost always the biggest single drop in the funnel. Consumption rate and basket size stand in for implementation, retention before lapse for persistence, and net price after promotion, discount and returns for gross-to-net.

The three failure modes transfer intact too: they forget, they had a poor experience, or they no longer feel they need it. The third is the hardest in both worlds, because the customer is not unhappy. They have simply stopped.

The commercially useful conclusion is the same sentence in both industries: one repeat rate across a portfolio is indefensible. It is a property of the category and the customer, not of the SKU.

I learned this from the wrong side of it. The business I have run for the last eleven years sells a considered, durable purchase, not a consumable. Durables do repeat, but the honest ratio is somewhere around ten to twenty per cent at best, and that single number reorganises the entire economics. ⭐ And a measurement trap worth naming. A repeat rate is meaningless without a window, and the convention is twelve months. But a durable good has a replacement cycle of three to seven years. So a twelve-month repeat rate on a durable is a truncated observation, not a low rate. It is exactly the censoring problem that makes you measure persistence over the full window rather than the first quarter. Ten to twenty per cent is what I observed inside a twelve-month cohort window. It is not the lifetime figure, and treating it as one would understate the business.

If most of your revenue is first-purchase, the acquisition cost has to be recovered on that first order. The break-even ceiling is first-order contribution plus the present value of whatever repeat you actually get, so at fifteen per cent repeat that is roughly 1.15 times first-order contribution. A consumable or subscription business runs at three to five times, because lifetime value recovers it later. Same margin, three to four times the bid capacity. ⚠ One caveat, because it is the thing people get wrong in the other direction. Bid capacity is average order value multiplied by contribution margin multiplied by that lifetime multiple, and repeat only drives the third term. Durables often have far higher order values, which is why mattress, furniture and luggage brands are among the highest CAC payers in direct-to-consumer. A high-value durable can and does outbid a cheap consumable. The constraint is not that you cannot pay. It is that you have one order to earn it back in, with no second chance. There is a compensating advantage nobody mentions, and it matters more since 2022. Paying a multiple of first-order contribution is a financed position: it needs working capital, it embeds retention risk in an assumption, and it is a lending decision against an unsecured future receivable. A business that recovers acquisition cost on order one has no payback gap and no exposure to a lifetime value that never arrives.

Which pushes a durable business toward the strategic answer I ran into: if repeat will not come from the same product, it has to come from the brand. You build a range of related products so the second purchase is a different item from the same trusted name, and the repeat rate you cannot get at SKU level you assemble at portfolio level. That reframes brand from a marketing indulgence into a financial mechanism, and it is the sort of thing a finance leader ought to be arguing for rather than defending against. It is also, for the record, a conclusion I reached around 2017 rather than one I am reverse-engineering now.

None of it excuses a weak product. In both models the product has to be right and fit for purpose, and no acquisition efficiency survives a product that disappoints. But when you live inside a low-repeat business for a decade you stop treating retention as a marketing metric and start treating it as the thing that decides whether the model works at all. It is also why I am unsentimental about acquisition-led plans: a business without a repeat mechanism has to win the same customer twice, and the forecast should say so out loud.

And here is a second transfer where the diagnosis is identical, though the remedy is not. At Parexel the constraint was clinical site utilisation: five sites, a fixed bed allocation paid for whether occupied or not, and contracts priced against planned utilisation rather than actual, so some work started below what it cost to deliver. The diagnosis travels perfectly. Any business carrying committed capacity mis-prices it the same way: against theoretical availability rather than realistic utilisation, and the overhead never comes back. That is room occupancy in a hotel, load factor in an airline, billable utilisation in a consultancy. The remedy splits, and knowing which is which is the point. Where capacity is sold forward under negotiated contracts, as in a CRO or a consultancy, you rebuild the rate card against realistic utilisation and gear bids to lift it. Where capacity is perishable, as in a seat or a room-night, you do close to the opposite: yield management, recovering fixed cost from the high-fare segment and taking anything above marginal cost for the rest. Price every seat at a full-absorption rate and you fly it empty. The trap in both is the same death spiral, raising rates because utilisation is low and losing the work that would have lifted it.

Where it does not transfer is the clinical layer. Epidemiology, lines of therapy, biomarker eligibility and reimbursement are their own discipline. What travels is the shape of the argument and the habit of building the number from the unit that causes it.

This is the subject of a follow-up piece, where I take the same cascade into consumer, subscription and technology businesses properly: recurring revenue, cohort retention, and why the SaaS metric set and the patient funnel are closer relatives than either side tends to admit.

A word on tooling, since people always ask. In life sciences the data usually comes from IQVIA, with Citeline or Evaluate for pipeline and analogues and Flatiron or Komodo for real-world evidence, while the model itself still tends to live in Excel with an enterprise planning layer such as Anaplan, Pigment or OneStream over it. In consumer and retail the equivalent stack is demand planning: o9, Blue Yonder, Kinaxis or RELEX, with Power BI or Tableau on top and a data-science platform such as Dataiku, Alteryx or Altair RapidMiner doing the preparation and the machine-learning layer. The names differ; the shape does not. And none of them decides whether your forecast is any good. An auditable cascade in a spreadsheet beats an elegant model nobody can open, every time.

Where I would start

On a mandate, this does not need a transformation programme.

Rebuild the forecast in patient units. Name the owner of every stage of the cascade, so each assumption has somebody accountable for it. Calibrate the treated pool against observed sales and find out which of the two is lying. Rank the assumptions by what actually moves the year, and put the two or three that matter in front of the board as arguable numbers rather than a single smooth line. Report a range with the drivers named, not a point.

That is days of work on most mandates, and it changes what the board is able to ask at the next meeting.

A board does not buy a forecast. It buys the ability to interrogate one, and a patient-based model gives it something to take apart in the room.

In brief

What is patient-based forecasting? Building the revenue number from the thing that causes it, rather than extrapolating the total. You start from the patient population and step down through diagnosis, treatment, label eligibility, access, share, persistence, adherence and dose, then apply a net price. Revenue is the answer at the bottom of the waterfall, not the input at the top.

Why not just use a trend line? For a stable product over a short horizon, often you can. The limitation shows when something structural moves, because a trend model tells you the pattern has broken without telling you which commercial driver broke it. It also cannot value a product that has not launched, which is most of what a pipeline decision is about.

What is the difference between adherence and persistence? Adherence is how completely a patient takes the prescribed dose, so it drives intensity. Persistence is how long they stay on therapy at all, and persistence is duration of therapy rather than an input to it. A patient can persist a full year at seventy per cent dose fidelity. Initiation is the third element and the one most often missed: whether a prescribed patient ever starts. The three have different fixes, which is the reason for separating them.

Why do patients not comply, and does the reason change the model? It changes the fix, which is what the model is for. Unintentional non-adherence is capacity: forgetting, age, cognitive decline, a complicated regimen. Persuasion does nothing; simplification, packaging and carer involvement do. Intentional non-adherence is a judgement, and it runs both ways. Punishing side effects in advanced disease make patients stop, which shows up as persistence and needs a tolerability answer. Feeling better makes patients stop too, and the relapse comes months later, which is the harder one to see and is the default risk in any asymptomatic chronic condition. Horne's necessity-concerns balance describes it: adherence tracks perceived need against perceived risk. So a single compliance rate across a portfolio is indefensible. It is a property of the disease and the patient, not of the molecule.

How much does the dosing regimen actually matter to a forecast? More than almost any other single assumption, and it is routinely entered as a round number. Published meta-analysis puts unadjusted adherence around 83 per cent for once-daily chronic medication and as low as 50 per cent for twice-daily, with timing adherence falling by roughly twenty-seven points. Assume ninety per cent where the reality is sixty-five and you have overstated volume by a quarter on that product. The regimen is also often a design decision, which is an argument for finance being in the room at target product profile stage rather than inheriting the assumption afterwards.

Why risk-adjusted NPV rather than NPV? Because a pipeline asset has a real chance of never reaching market. The mechanics matter: you weight each cash flow by the probability that it occurs, conditional on every prior development transition, rather than multiplying a finished NPV by a single number. And the discount rate has to be consistent with it, or the same risk gets counted twice.

What does AI change? The genuine new capability is estimating a biomarker-eligible population from claims and records where the population was never fully tested, plus extracting structured epidemiology from unstructured literature. It also removes the cost constraint that used to ration this to the biggest decisions only. The limits are real: explainability to a board, model drift, and data protection and sparsity in exactly the rare-disease settings where it would help most. Model logic and aggregate inputs only, version control, independent recalculation, and a named human owner for every assumption. The judgement stays human.

Sources

© Jatinder Purewal 2026. All rights reserved.

If your revenue line cannot say which assumption is carrying it, that is usually days of work rather than a programme. I take those conversations directly: get in touch.

Discuss this piece on LinkedIn: linkedin.com/in/jatinderpurewal · More Insights