Decomposing Sales Forecast Error: Estimation, Structural and Incentive Sources
August 3, 2026
ABSTRACT
Forecast error in revenue organisations is treated as a discipline problem to be solved by more frequent review. The forecasting research literature treats judgmental forecasting as a well-characterised process with known biases and known corrections, and its findings apply to pipeline forecasting almost without translation. This paper maps those findings onto sales forecasting practice, sets out an error-decomposition method that separates the three distinct sources of error, and identifies which of them additional review meetings can and cannot address.
A sales forecast is a judgmental forecast produced by people whose compensation depends on the outcome being forecast, aggregated through a model nobody has validated, and reviewed by a process that mostly asks whether individual numbers feel right.
Stated that way, it is unsurprising that accuracy is poor. It is also unnecessary, because judgmental forecasting is one of the better-studied problems in applied research, and the findings transfer with very little adaptation.
1. What the forecasting literature establishes
Armstrong, Green and Graefe [1] synthesise a large body of forecasting evidence into a single unifying principle: be conservative. Conservatism here has a specific meaning — a forecast should be conservative with respect to cumulative knowledge about the situation and about forecasting itself. In practice that means anchoring on established base rates, damping trends rather than extrapolating them, combining independent forecasts, and not over-reacting to recent observations.
Each of those is routinely violated in pipeline review. Base rates — historical win rates by stage, segment and source — are frequently not consulted at all. Trends are extrapolated from a single quarter. Forecasts are produced by one hierarchy rather than combined from independent ones. And recent movement dominates the discussion, because the discussion is organised around what changed this week.
2. The judgement layer and its known biases
Underneath the process sits an individual estimating a probability. Tversky and Kahneman [2] characterised the heuristics people use for exactly this task and the systematic errors those heuristics produce. Three are directly relevant.
The general findings are from the judgement literature; the pipeline expressions are the application proposed here.
| Heuristic | General effect | How it appears in a deal review |
|---|---|---|
| Representativeness | Judging likelihood by similarity to a type, ignoring base rates | "This looks exactly like the deal we won last quarter" — base rate not consulted |
| Availability | Judging frequency by ease of recall | Recent losses depress estimates across unrelated deals |
| Anchoring and adjustment | Insufficient adjustment away from an initial value | The first-entered close date and amount persist through evidence that should move them |
Source: Heuristics from Tversky & Kahneman (1974); application proposed here
Why this is not solved by asking people to try harder
3. Three sources of error, which are usually conflated
Most forecast post-mortems produce a single number, and a single number cannot indicate what to change. There are three distinct sources and they have different remedies.
Estimation error. The individual judgement about a specific deal was wrong. This is the source review meetings target, and it is usually the smallest of the three in a mature organisation.
Structural error. The aggregation model is wrong — stage probabilities that no longer match observed conversion, a weighting scheme never re-fitted, categories whose definitions have drifted. This error is systematic, invisible at the deal level, and survives any amount of review because the review operates below it.
Incentive error. The forecast is used for something other than prediction — quota setting, resource allocation, performance assessment — so its producers optimise for that use. Ridgway [3] and Oyer [4] together establish both the general mechanism and its observable effect on the timing of business activity.
The practical significance: adding review cadence targets source one only. Where sources two and three dominate — which is common — it adds cost and changes nothing.
4. Method: decomposing forecast error
Run this over at least eight closed periods. Fewer than eight cannot separate bias from noise.
Compute mean error and mean absolute error separately, per period. Mean error is bias — a persistent direction. Mean absolute error is noise — dispersion irrespective of direction. Reporting only one is the most common measurement failure in forecast review, and the two demand opposite responses.
Test the structural layer. For each stage, compare the assigned probability against the observed historical conversion of opportunities that entered that stage. Where assigned and observed diverge, the error is structural and no amount of deal review will find it.
Test the estimation layer. Take the earliest submitted forecast for each deal and compare it against the outcome, grouped by forecaster. Consistent directional error by individual is estimation bias and is correctable with a personal adjustment factor.
Test the incentive layer. Plot the timing of forecast category changes against period boundaries. Systematic movement in the last week of a period reflects the use of the forecast rather than information about the deals.
Attribute the total. Assign each period's error across the three sources. The largest source is the one to work on, and it is frequently not the one currently receiving attention.
5. Corrections, in order of yield
Re-fit stage probabilities against observed conversion, then re-fit them on a schedule. This is arithmetic on data already held and it is the highest-yield step in most organisations.
Introduce a base-rate comparison into deal review. Before the qualitative discussion, state the historical win rate for opportunities of this segment, source and stage. This directly addresses the representativeness effect.
Combine independent forecasts. A model-derived forecast produced alongside the judgmental roll-up, with divergences examined rather than reconciled away, is the single most consistently supported recommendation in the forecasting literature [1].
Damp trend extrapolation. A quarter of improved conversion is weak evidence of a new rate. Conservatism with respect to cumulative knowledge means weighting the longer history more heavily than the recent movement.
Separate the forecast from the assessment. Where the same number sets quota and evaluates performance, incentive error is structural. Producing an unassessed forecast alongside the committed one is a low-cost way to measure how large that effect is.
6. Limits of this argument
The cited work establishes a general conservatism principle supported across forecasting domains [1] and the existence of systematic biases in judgement under uncertainty [2]. Neither was conducted on B2B sales pipelines specifically, and the application here is an argued transfer rather than a demonstrated one.
The transfer rests on a structural similarity that is easy to check: pipeline forecasting is judgmental estimation under uncertainty, aggregated through a model, by parties with an interest in the result. Where that description fits, the findings should be expected to apply. The decomposition method in section 4 does not depend on the transfer being correct — it measures the local situation directly, which is the point.
Get Decomposing Sales Forecast Error: Estimation, Structural and Incentive Sources as a print-ready PDF.
Includes the full reference list. One form unlocks every paper and template on the site.
COMMON QUESTIONS
- Why is sales forecast accuracy so poor?
- Three distinct sources are usually conflated: estimation error in individual deal judgement, structural error in the aggregation model, and incentive-induced error from how the forecast is used. Each has a different remedy, and review meetings address only the first.
- Does more frequent forecast review improve accuracy?
- It addresses estimation error at the deal level and does nothing for structural or incentive error. Where those dominate, adding review cadence adds cost without changing the outcome.
- What does the forecasting research recommend?
- Armstrong, Green and Graefe (2015) synthesise the evidence into a conservatism principle: forecasts should be conservative with respect to cumulative knowledge — use established base rates, avoid over-reacting to recent movement, and combine methods rather than relying on a single judgement.
- How do you separate forecast bias from forecast noise?
- Compute mean error and mean absolute error separately over at least eight periods. Bias is systematic and correctable by adjustment; noise is dispersion and is not. Reporting only one of the two is the most common measurement error in forecast review.
KEY TERMS
Time in Stage
How long opportunities remain in each pipeline stage. The diagnostic that localises where deals actually stall.
Forecast Accuracy
How closely forecasts match actual results, measured consistently over time. The metric that separates disciplined forecasting from lucky forecasting.
Close Date
The date an opportunity is expected to be signed. The single most consequential field in a pipeline record.
Sales Forecast Category
A judgement label applied to an opportunity — commit, best case, pipeline — expressing confidence independently of stage.
Pipeline Coverage
The ratio of open pipeline value to the target for a period. Answers whether there is enough in play to hit the number at historical conversion rates.
Weighted Pipeline
Pipeline value multiplied by a probability assigned to each stage, producing an expected value rather than a gross total.
WHERE THIS HAS BEEN APPLIED
Client work and research from RevOps HQ, our consulting practice.
REFERENCES
- [1]Armstrong, J. S., Green, K. C. & Graefe, A. (2015) Golden Rule of Forecasting: Be Conservative Journal of Business Research, 68(8), 1717–1731 link
- [2]Tversky, A. & Kahneman, D. (1974) Judgment under Uncertainty: Heuristics and Biases Science, 185(4157), 1124–1131 link
- [3]Ridgway, V. F. (1956) Dysfunctional Consequences of Performance Measurements Administrative Science Quarterly, 1(2), 240–247 link
- [4]Oyer, P. (1998) Fiscal Year Ends and Nonlinear Incentive Contracts: The Effect on Business Seasonality The Quarterly Journal of Economics, 113(1), 149–185 link
Put this into practice: RevOps 101: Revenue Operations Foundations
The course walks through the same material with the interactive audit tools.
See the course