Baseline, judgement, and the most misused word in forecasting
Neither the model nor the manager is "the biased one". Bias is a property of a process, established against actuals, and the back-test decides whose changes are earning their keep.
I heard a familiar claim again on a podcast recently, stated with complete confidence: an AI model produces a "biased" forecast, and human expertise is what makes it "unbiased". It is a natural thing to say. Most of the room will have nodded. And it is wrong in a way that matters, because the correction is not a piece of pedantry: it is the difference between a forecasting process you can govern and one you can only argue about. What people usually mean by the claim is that models miss context. That is true, and it is rule one below. The error is in the word, not the instinct.
A note on scope: this piece is deliberately not specific to any sector. The governing principles transfer across sectors, although the metrics and the implementation differ. The concepts are old and none of them are mine: the research base is decades deep and I cite it at the foot. What is mine is the experience of running forecast processes against actuals since 1999, when I ran group budgeting, forecasting and long-range planning cycles at IMS Health's global headquarters, then across divisional, European and group finance seats, and currently on my own profit and loss account, and the operating habits I have watched separate the processes that improve from the ones that just continue. Where something is my judgement rather than evidence, I say so.
The word has a definition, and it is not "came from a machine"
In forecasting, bias has a precise meaning: a systematic tendency to run high or low, a property of the forecasting process rather than of any single number it produces, established by scoring signed errors against actuals across a run of comparable forecasts. One number, one period, is not enough to establish it. A forecast that came in 8% over is not "biased"; it is inaccurate, once. Bias is the pattern: over, and over again, in the same direction.
Which leads to the part the claim gets exactly backwards. You cannot establish whether a forecast is biased on the day it is made, and you certainly cannot establish it from where the forecast came from. A machine-generated number and a human-adjusted number are both capable of being biased, and both capable of being unbiased on the chosen measure. There is one honest exception: where a method is structurally certain to run one way, a flat average against a steadily growing market, the bias is predictable before it is measured, and the answer is to fix the method. For everything else, the only way to find out is the unglamorous one: score everything against actuals, for long enough for a pattern to show.
Three different things are being called one thing
The confident podcast sentence collapses three separate concepts into one word, and most forecast arguments inside companies do the same.
Accuracy starts from how close each forecast landed to its actual. Strictly, the per-period gap is the forecast error; accuracy is that error aggregated across forecasts under a named measure, and different measures can rank the same forecasts differently.
Bias is which direction it tends to miss, consistently: a pattern, measured across periods.
Judgemental adjustment is neither of those. It is an intervention: the act of a human changing a number that a process produced. An adjustment is not good or bad by nature. It is an input whose effect on accuracy and bias can be measured, and the research says its effects are anything but uniform.
Keep the three apart and something useful becomes visible: an adjustment can improve accuracy while creating bias. Using the convention that a positive error means the forecast ran above actual, suppose the baseline's errors alternate, plus 20 then minus 20: the average miss is 20 and the bias is zero. Now adjust so the errors become plus 5 every period: the average miss falls to 5 and a bias of plus 5 has appeared from nowhere. Accuracy improved, bias created. The reverse trade exists too. A point forecast tuned to minimise the average absolute miss targets the median; one tuned to minimise squared error targets the mean. Under right-skewed demand those are different numbers. A median-targeting forecast will sit below the mean of what actually happens, period after period, and so will score as biased low on a signed-error measure while doing exactly the job it was tuned for. That is not a paradox. It is a reminder that bias is only meaningful against a stated target and a stated measure. Pick both to fit the decision: a stockout costs a sale and sometimes a customer, an overstock costs cash and eventually a write-off, and those two mistakes are not the same size. And an aggregate bias of zero can hide severe conditional bias, in promotional periods, at high volumes, in one region. Collapse the three concepts into one word and every conversation about the forecast becomes a conversation about whose number wins, which is politics, not process.
The vocabulary that ends the argument
The practical fix is not a lecture on statistics. It is naming things properly in the planning pack, because the vocabulary does the governing for you.
Call the machine's number the baseline forecast: the statistical or AI-generated starting point, produced from history and available signals, on a documented method. Call the human layer what it is: a judgementally adjusted forecast. Call the thing the board sees the decision-ready forecast, and let the pack show the bridge from baseline to decision-ready as a visible set of adjustments, each with an owner and a stated reason.
None of these labels accuses anybody of anything. That is the point. "Baseline plus adjustments" replaces "your number versus my number", and it makes the next section possible, because you cannot score what you never separated. And keep the forecast separate from the target, the plan and the commitment. A number that has been moved because someone needs it to be that number is not an adjustment, it is a negotiation, and scoring it as a forecast corrupts the back-test.
What the evidence actually says about adjustments
The judgemental-adjustment literature has a landmark study by Fildes, Goodwin, Lawrence and Nikolopoulos, published in 2009: more than sixty thousand real forecasts and outcomes from four supply-chain companies, comparing the statistical baseline with the judgementally adjusted number that replaced it.
Three findings from that work should shape how any finance team runs its process. Adjustments improved accuracy on average in three of the four companies, so the caricature of the meddling manager who should leave the model alone is not supported: judgement added accuracy in three of the four processes studied. Larger adjustments tended to be more effective than small ones, as an association rather than an instruction to inflate anything. Small adjustments frequently made the forecast worse; my interpretation is that they tend to encode noise, ownership signalling and the urge to touch the number, rather than information. And upward adjustments were much less likely to help than downward ones, and more often wrong in direction, a pattern the authors interpreted as optimism bias. The same lead author's later pooled evidence is colder and more useful. Across roughly 147,000 forecasts from six studies, re-analysed on a common framework, adjustments improved bias and accuracy for only just over half of items. Positive adjustments were confirmed as more likely to make things worse, while negative ones typically improved matters, particularly large ones, which puts a sign on the "larger is better" association above. And the evidence that forecasters were using information the algorithm did not have was weak: they appeared to be responding to irrelevant cues, or cues of little diagnostic value.
That reads like a case against the human layer. It is a case against an unmanaged one, and it is the best argument in the literature for the habits below. An adjustment that has to be named, owned and justified in one sentence before the outcome is known is much harder to make on an irrelevant cue. The same paper points at the remedy: a debiasing procedure applied to adjusted forecasts improved performance. Adjustment is not the problem. Unscored adjustment is.
That finding travels to my own sector. In the same journal issue, Franses and Legerstee examined expert adjustments to pharmaceutical SKU forecasts across thirty-seven countries: experts adjusted frequently and mostly upwards, and their adjustments were largely predictable, with a forecaster's own previous adjustment carrying roughly three times the weight of the model's past errors. The persistence suggests adjustment was being driven at least as much by the forecaster's own previous interventions as by evidence about where the model had been wrong.
Read those findings together and the message is not "trust the model" or "trust the manager". It is: adjust when you know something the baseline has not been given, make it count, and subject the adjustments that make the number bigger to markedly more scrutiny.
Two kinds of bias, and only one of them is a modelling problem
The word covers two different machines, and they take different spanners.
A baseline's bias is usually an artefact of what it was trained on. The cleanest example in demand forecasting is censored history: the model learns from what was sold, not from what was wanted. Every stockout teaches it that demand was lower than it really was, so it under-forecasts precisely the lines that keep running out, and the shortage renews itself. Nothing in the model is broken; it answered the question it was given, on the data it was given. Structural breaks do the same job from the other end, a model fitted through a growth period carrying that slope into a market that has stopped growing. Both are fixed by changing the data or the specification. Neither is fixed by argument.
Human bias in a corporate forecast often has a different engine, and this is where I part company slightly with the adjustment research. Those studies measure the behaviour carefully and explain it mostly in cognitive terms: anchoring, optimism, the urge to touch the number. In my experience a large share of the bias in a company's numbers is not cognitive at all. It is structural, and from the forecaster's side of the table it is entirely rational.
Take the sales forecast that becomes next year's target, with a bonus attached to beating it. You are asking someone to set the bar they will be measured against, and trusting them to set it honestly. A number that comes in conservative is not a mistake. It is the rational answer to the incentive in front of them. Management accounting has had a name for this since 1970, budgetary slack: revenues understated and costs overstated by the people closest to both, on a scale one early study put at a fifth to a quarter of a division's operating costs. The general form is the principal-agent problem. The person with the information has an interest that does not match the person relying on it.
Notice the direction of travel. The supply-chain studies find optimism, adjustments biased upward. The budget cycle frequently produces the opposite, forecasts nudged down and held down by people protecting the number they will be judged on. Same word, opposite signs, because the incentives point in opposite directions, and that is the tell that the cause is structural rather than psychological. I flag the reading as my judgement: the forecasting research and the budgeting research have each studied this for decades and are rarely put in the same room. The exception worth knowing is Oliva and Watson's case study, which splits functional bias into the intentional kind, driven by misaligned incentives and where power sits, and the unintentional kind that comes from blind spots, and then shows a consensus process managing both.
This is the part of the job a finance leader is actually hired for, and it is not being difficult in meetings. Challenge that works is structural rather than personal:
- Separate the forecast from the target, and score them separately. A forecast is a best estimate; a target is a commitment. The moment one number does both jobs, somebody is being paid to bias it.
- Publish the back-test. Slack survives a conversation. It rarely survives a run of scored forecasts with an owner's name against it, which is why measurement beats argument.
- Ask what would have to be true. Sandbagging hides in the aggregate. It seldom survives a walk through the assumptions underneath, because the slack has to be sitting somewhere specific: a conversion rate, a launch date, a price.
- Reset the anchor when the anchor is the problem. Where the issue is inertia rather than incentive, a cost base nobody has questioned in five years, zero-based budgeting makes every line justify itself from nothing. It is expensive, and run every year forever it becomes theatre, but as a periodic reset it does something no adjustment ever will.
None of that requires accusing anyone of anything, which is the same principle as the vocabulary. Change the structure and the behaviour changes with it.
Forecast value added: the governing test
The test that turns all of this into governance has a name: forecast value added, a discipline popularised by Mike Gilliland. The question it asks is brutally simple. Did each step in your process beat the step it replaced, measured against actuals?
Run the chain: a naive baseline, last period carried forward, or a seasonal naive where the series demands it, costs almost nothing and sets the zero point; which naive comparator you use is a choice, so fix it and disclose it. A statistical or AI baseline has to beat naive on the horizon and measure that matter for the decision, or it is expensive decoration. The comparison only means anything if the horizon, the information date, the aggregation level and the error measure are held constant across steps. The judgementally adjusted forecast has to beat the baseline it overwrote, or the adjustment process, whatever it feels like from inside the meeting, is subtracting value. Every step then lands in one of three findings: it adds value, it subtracts value, or the difference is not yet material or statistically credible, and the third verdict is a real answer, not a failure to decide. Where value is subtracted, that is a finding about the process, not a verdict on the people.
FVA is not uncontested, and a piece that recommends it should say so. The published critique runs roughly as follows: that a step which happens to score well may have been lucky rather than good; that accuracy is not the same thing as money; and that a point-forecast view of the world ignores the uncertainty the decision actually faces. The first charge is the one worth answering, and habit 2 below is the answer. A reason written at the moment of adjustment, before the outcome is known, is a claim staked in advance, and it is what separates an adjustment that added value from one that got lucky. On the second charge I would concede rather than argue. FVA scores the process, not the payoff. So finish the chain with the economic test: a step that improves accuracy still has to improve a decision by more than the step costs to run, or it is a well-governed waste of time.
This is the correction to the podcast claim in operational form. The question was never which side is biased, the machine or the human. The question is whose changes are adding value, and the only honest answer comes from the back-test.
Where I learned to score everything
I learned this discipline in the years when the baseline was a spreadsheet, not a model. It started at group level: my first corporate seat, in 1999, was the budgeting, forecasting and long-range planning cycle at IMS Health's global headquarters, where the consolidated numbers arrived from the businesses and the reconciliation between what had been forecast and what had actually happened was somebody's job every quarter. Mine. At Chiron, in the year I ran European and international FP&A for the biopharmaceutical business, I rebuilt the in-year forecasting and the reporting that scored it: budgeting and forecast templates, and automated bridge reporting at profit-centre and cost-centre level, across the profit-centre and cost-centre owners of a region spanning EMEA and Asia Pacific, where actuals uploaded and the variance against forecast and budget fell out on its own. Forecast error, the absolute gap between the in-year forecast and actual results, fell from roughly 15% to roughly 4% on the same internal reporting basis. A forecast you can trust in-month is what lets a business commit stock, cash and headcount earlier in the cycle rather than holding a buffer against its own planning. The principal mechanism, in my assessment, was not a cleverer forecast. It was measurement with consequences: every forecast compared with what actually happened, every material miss decomposed to its driver, and the next cycle's assumptions adjusted by people who knew their last assumption had been scored. The improvement was the by-product. The habit was the asset.
And it travelled with me. At Mayne Pharma I owned the budgeting and long-range planning cycles at group level for a global generics business, and rebuilt the monthly reporting that went to senior leadership. In the Early Phase unit at Parexel I drove it hardest: management reporting rebuilt around the drivers that actually moved the numbers, with bridge and waterfall analysis running into the forecast and the commentary, devolved into the team so every analyst could see their own picture rather than wait for mine. At Shionogi it was the month-end close and forecast cycle into a Japanese parent. Different companies, different systems, one question every month: what did we say would happen, what happened, and who owns the difference.
Since then the tooling has changed completely; the habit has not. On my own company's numbers today, the baseline for a demand or cash forecast comes from an AI-assisted model: I specify it, Claude and ChatGPT write and check the model code, and the model runs locally against my own sales history and order data, with the evaluation out of sample; the code goes to the models, the data does not, and client work runs under the client's own data-governance rules. The model file is versioned, the evaluation window is held out, and drift gets rechecked when the market moves. The result is fast and consistent, and across the monthly cycles scored so far the back-test shows no systematic upward tilt. The judgement layer is still mine, because the model has not been told that a supplier's lead time has quietly moved, that a delisting was agreed last week, or that a price move lands in March. And the scoring is still the point: baseline and adjustments alike are back-tested against actuals, monthly, using the forecast as it stood at the time, and some months the uncomfortable finding is that the model's number was better before I improved it. I should be honest about scale: this is the single-owner version of the discipline, scored monthly on a live P&L; it demonstrates the habit, and the multi-owner governance below is what I have run at corporate scale. That continuity is the reason I trust the discipline more than I trust any particular forecast, though I flag that conviction as judgement, built on my own runs of data rather than a controlled study.
This is the fairest way I know to describe what AI changes in forecasting, and what it does not. In my own implementation the baseline got materially better and much cheaper to produce, and that experience generalises with a caveat: where data are sparse or the business keeps breaking structurally, a simple benchmark still takes some beating. The need to separate, own and score the human layer does not move at all. If anything it sharpens, because a better baseline raises the bar every adjustment has to clear. The freshest evidence points the same way: a 2025 study put 123 human forecasters against five large language models, and the models did not consistently win, while both humans and models got materially worse in promotional periods, which is conditional bias by another name. I wrote about the modelling side of this in the patient-based forecasting piece; this article is the governance that has to sit around any such model.
The operating burden, which is the actual answer
None of this works as a memo. It works as a set of operating habits, and they are unglamorous:
1. An owner per assumption. Every material adjustment has a name attached, not a department.
2. A written rationale at the moment of adjustment. One sentence, captured when the change is made, because a reason reconstructed at review time is a story, not a rationale.
3. A regular back-test, published internally. Baseline versus adjusted versus actuals, each owner scored against their own baseline on the same items, with the number of adjustments shown so the reader can see whether the sample supports the finding. Never a league table across owners. Per-owner samples are rarely large enough to support a ranking, which is the third finding applied to people rather than to steps; and a published ranking turns a diagnostic into a threat, at which point the rational response is to stop adjusting at all, and the process loses the information it exists to capture. Cadence follows volume: quarterly for a handful of material assumptions, faster where the volumes support it.
4. The uncomfortable conversation. When someone's adjustments have subtracted value across enough scored cases for the finding to be real rather than noise, the process has to be allowed to say so. The trigger is a conversation, not a verdict, and the first question in it is whether the score is signal at all. If it is, the question that follows is not "why are you biased". It is "what are you seeing that the baseline has not been given?" Sometimes the answer names a signal worth building into the model. Sometimes there is no answer, and the honest outcome is that the adjustment stops and the person contributes the signal instead of the number.
5. A debiasing pass on the adjusted set. The pooled evidence's own remedy: where the back-test shows a persistent tilt in a layer of adjustments, apply a standing correction to that layer until the tilt closes at source. Mechanical, cheap, and it turns a known bias from a debate into an entry.
And where there is no statistical baseline at all, because the forecast is a spreadsheet owned by whoever argues best, the first move is not a model. It is the same vocabulary: name that spreadsheet the baseline, freeze it at a date, and start scoring it against actuals. The case for anything better builds itself from the scores.
Compressed to a paragraph, that is where I stand:
"I see AI as producing an excellent baseline forecast from historical data and available signals. Management then overlays operational knowledge, commercial insight, strategic initiatives and one-off events that are not present in the data. The objective is not to remove bias; it is to improve forecast quality and decision usefulness. We then evaluate the process over time by measuring both forecast accuracy and forecast bias against actual results."
The forecast argument inside most companies is about whose number wins. The fix is vocabulary and a back-test: separate the baseline from the adjustments, then let actuals decide who is earning their keep.
In brief
Is an AI forecast biased? Not by origin. Bias is a property of a forecasting process, a systematic directional miss established by scoring against actuals across comparable forecasts, and a machine baseline and a human-adjusted number are both capable of it. The question is not which side is biased but whose changes add value.
Should managers adjust a statistical or AI baseline? Yes, when they know something the baseline has not been given: across more than sixty thousand real forecasts and outcomes, adjustments improved accuracy on average in three of the four companies studied. But small tweaks frequently made forecasts worse, and upward adjustments helped least; pooled evidence across roughly 147,000 forecasts finds adjustments helped for only just over half of items, with the gains concentrated in large downward ones. So adjust rarely, make it count, and score every adjustment against actuals.
Why do forecasts get biased in the first place? Two different reasons that need two different fixes. A model's bias is usually an artefact of its data: censored history teaches it to under-forecast the lines that keep selling out, and a structural break leaves it carrying yesterday's slope. A person's bias in a corporate forecast is often an artefact of the incentive: ask someone to forecast the number they will be bonused against and conservatism is rational, which management accounting has called budgetary slack since 1970. The first is fixed by changing the data or the specification; the second by separating the forecast from the target and scoring both.
What is forecast value added? The governing test of a forecast process: each step, from naive baseline to statistical or AI baseline to judgementally adjusted forecast, must beat the step it replaced, measured against actuals. A step that subtracts value is a process finding, and the honest response is to change the process.
What actually improves forecast accuracy over time? Measurement with consequences creates the feedback that improvement needs. In my own experience, from cutting forecast error, the absolute gap between the in-year forecast and actuals on a consistent internal basis, from roughly 15% to 4% in the year I ran European and international FP&A at a pharma business, to back-testing AI baselines on my own company's numbers today, the improvement follows the scoring: owners named per assumption, rationales written at the moment of adjustment, and a regular internally published back-test.
If your forecast pack cannot show the bridge from the baseline to the number the board saw, with owners on each adjustment, that is usually the first thing worth fixing. I take those conversations directly: get in touch.
Discuss this piece on LinkedIn: linkedin.com/in/jatinderpurewal · More Insights
References
- Fildes, R., Goodwin, P., Lawrence, M. and Nikolopoulos, K. (2009), "Effective forecasting and judgmental adjustments: an empirical evaluation and strategies for improvement in supply-chain planning", International Journal of Forecasting, 25(1), 3-23 - the 60,000+ forecasts-and-outcomes study across four companies.
- Franses, P.H. and Legerstee, R. (2009), "Properties of expert adjustments on model-based SKU-level forecasts", International Journal of Forecasting, 25(1), 35-47 - pharmaceutical expert adjustments across thirty-seven countries: frequent, mostly upward, and largely predictable.
- Fildes, R., Goodwin, P. and De Baets, S. (2025), "Forecast value added in demand planning", International Journal of Forecasting, 41(2), 649-669 (published online 2024) - pooled evidence across roughly 147,000 forecasts from six studies: adjustments helped for just over half of items, positive adjustments confirmed harmful, large negative ones most effective, and a debiasing procedure on adjusted forecasts improved performance.
- Doherty, C. (2024), "A Critical Evaluation of the Assumptions of Forecast Value Added", Foresight: The International Journal of Applied Forecasting, Issue 73 (Q2 2024), 13-23, with commentaries - the published critique of FVA answered in the text.
- Abolghasemi, M., Ganbold, O. and Rotaru, K. (2025), "Humans vs. large language models: judgmental forecasting in an era of advanced AI", International Journal of Forecasting, 41(2), 631-648 - 123 human forecasters against five LLMs; neither side consistently superior, both worse in promotional periods.
- Hyndman, R.J. and Athanasopoulos, G., Forecasting: Principles and Practice (3rd edition), chapter 6, "Judgmental forecasts" - otexts.com/fpp3.
- Gilliland, M. (2010), The Business Forecasting Deal, and the SAS white paper "Forecast Value Added Analysis: Step-by-Step" - the FVA framework for scoring each step of a forecast process against the step it replaced.
- Oliva, R. and Watson, N. (2009), "Managing functional biases in organizational forecasts: a case study of consensus forecasting in supply chain planning", Production and Operations Management, 18(2), 138-151 - functional bias split into the intentional kind, from misaligned incentives and the disposition of power, and the unintentional kind, from informational and procedural blind spots.
- Schiff, M. and Lewin, A.Y. (1970), "The impact of people on budgets", The Accounting Review, 45(2), 259-268 - the origin of budgetary slack: revenues understated and costs overstated by managers influencing the standard they will be judged against.
- Institute of Business Forecasting glossary - standard terminology for baseline, judgemental adjustment, forecast accuracy and bias.
The views here are my own. They do not represent the position of any current, former or future employer or client, and nothing here draws on confidential information.
© Jatinder Purewal 2026. All rights reserved.