Forecast Bias: The Formula, the Sign Trap, and When It Beats Chasing Accuracy

The Supply Chain GuysHonest supply chain judgment from practitioners

The forecast bias formula is the sum of actuals minus the sum of forecast, divided by the sum of forecast, over the same periods and the same lag: (ΣA − ΣF) ÷ ΣF × 100%. A positive result means you under-forecast and sold more than you planned, a negative result means you over-forecast, and anything that stays the same sign for several periods in a row is bias rather than noise. Bias matters more than the size of your error, because safety stock is designed to absorb random misses and is defenceless against directional ones.

What is the forecast bias formula?

Add up what actually happened, subtract the forecast, and divide by the forecast. That is it. The arithmetic is trivial and the interesting decisions all sit around it: which periods, which lag, which level of aggregation, and which sign convention. Bias measured at one lag on one aggregated total is a number that will comfort you. Bias measured at your operational lag, per product and per customer, is a number that will start arguments worth having.

The mean signed error is the same measurement without the percentage, (ΣA − ΣF) ÷ n, and it is often more useful when you are looking at a single item in units.

Which sign convention should you use?

Four are in common use and they disagree with each other. Getting two of them into the same meeting is how a demand review spends twenty minutes discovering that everybody agrees.

ConventionFormulaA positive number meansWhere we usually meet it
Deviation over forecast (our house standard)(ΣA − ΣF) ÷ ΣFWe under-forecast, actuals beat the planSales and finance reporting, S&OP attainment views
Error over actuals(ΣF − ΣA) ÷ ΣAWe over-forecastDemand planning packs, most planning software
Mean signed error(ΣA − ΣF) ÷ nWe under-forecast, expressed in unitsItem-level review, safety stock sizing
Forecast attainment indexΣA ÷ ΣFAbove 1.0 means actuals beat the planCommercial reviews, supplier scorecards
Four sign conventions in common use. Ours is the first row. Label whichever one you use, on the report itself.

Pick one, write it at the top of the report, and never let the other three appear in the same deck without a label.

Ours is deviation over forecast: (ΣA − ΣF) ÷ ΣF, where a positive number means we sold more than we planned. Write that sentence on the report itself, in words, not just the symbols. The failure we run into most often is not a wrong formula. It is a page that prints a signed percentage and never says which way the sign points, so half the room reads it backwards and nobody finds out until somebody sizes a buffer on it.

We standardise this way because it is the number the commercial side of the business already speaks. A demand review that sits next to a sales review has to agree what "we were 6% out" means before it can agree on anything else, and deviation against plan is the version both rooms turn up already using. It is also the convention this site has run on since the 2016 piece this one follows from, so the back catalogue and the new work read the same way round.

One thing does move, and it is worth being straight about because it is our own back catalogue. Mean absolute error is often written as Σ|F − A| ÷ Forecast, and that is how the older pages here wrote it. Dividing by the forecast measures error against the plan you made, which flatters a business that plans low, punishes one that plans high, and moves the denominator every time the forecast moves. Our house standard divides by actuals, Σ|F − A| ÷ ΣA, as a percentage of what we actually sold. That gives you WAPE, it anchors the number to what really happened, and it stays comparable across periods and across items. Both versions are in live use across the industry, and you will meet the forecast-denominator one in plenty of places, our own earlier writing included (the forecast accuracy chapter of Hyndman and Athanasopoulos sets out the scale-free error measures this family sits in). We have standardised on actuals from here on and left the older pieces as they were published, rather than quietly rewriting the back catalogue to match a decision taken later.

Worked example: the same six months, two different answers

Worked example built to show the mechanics, not a client data set.

MonthForecastActualError (A − F)Absolute error
Jan1,000940−6060
Feb1,1001,010−9090
Mar950900−5050
Apr1,2001,150−5050
May1,050980−7070
Jun1,1501,080−7070
Total6,4506,060−390390
Six months of directional error. Illustrative figures, not client data.

House standard, deviation over forecast: −390 ÷ 6,450 = −6.0%. Negative, so we over-forecast in every month of the half and carried the stock for it. Run the same six months as error over actuals and the number is 390 ÷ 6,060 = +6.4%. Same spreadsheet, same half, opposite sign and a different magnitude. Neither is wrong. Both in one room with neither one labelled is how a business spends a quarter arguing about a direction it had already agreed on.

Is it bias or is it noise?

Compare the size of the signed total to the size of the absolute total. Call it the bias ratio: |Σ error| ÷ Σ |error|. It runs from 0 to 1. Near zero means your misses cancel out and you are looking at noise, which is what safety stock exists to cover. Near one means almost every unit of error points the same way and the buffer will never catch up.

In the table above, |−390| ÷ 390 = 1.00. Every month landed on the same side of the forecast. That is as pure as bias gets.

Now the same MAE with the error scattered:

MonthForecastActualError (A − F)Absolute error
Jan1,0001,060+6060
Feb1,1001,010−9090
Mar9501,000+5050
Apr1,2001,150−5050
May1,0501,120+7070
Jun1,1501,080−7070
Total6,4506,420−30390
Identical absolute error to the table above, scattered rather than directional. Illustrative figures, not client data.

Identical mean absolute error of 65 units. Bias ratio |−30| ÷ 390 = 0.08. The first business needs a governance conversation. The second one needs a bigger buffer or a better model, and no amount of shouting at planners will move it. Same error size, completely different problem, and a single accuracy KPI cannot tell them apart.

This is the same idea as the classic tracking signal (cumulative signed error divided by mean absolute deviation), scaled to sit between 0 and 1 so it reads the same on any item. Use whichever your system already produces. Just make sure one of them is on the page.

Bias versus noise: the same average error size, two different problems Two line charts over six months. On the left, the forecast line sits above the actuals line in every single month, giving a bias ratio of 1.00. On the right, the same average gap alternates above and below the actuals line, giving a bias ratio of 0.08. Both have an identical mean absolute error of 65 units. Directional error (bias) Bias ratio 1.00 · every month high JanFebMar AprMayJun Forecast above actuals in all six months. The buffer never catches up. Random error (noise) Bias ratio 0.08 · misses cancel out JanFebMar AprMayJun Same average gap, alternating direction. This is what safety stock is for. Actuals Forecast
Both panels carry an identical mean absolute error of 65 units. Illustrative figures, not client data.

What does low-grade forecast bias actually cost?

Enough to eat a year of margin on a product line, and it can do that while the headline bias number is pointing the other way.

We watched this happen at a business whose aggregate bias read as consistent under-forecasting. On the face of it that is the safer direction to be wrong in. It usually means you are leaving a bit of service on the table rather than money. What the aggregate was really reporting was one product line with a production capacity constraint, where genuine demand ran well ahead of anything anyone was willing to put in the forecast.

The sales team knew about the constraint. They also had a target. Since the SKU in real demand could not be made in the volume that target needed, the forecast went onto the SKUs that could. The number was not wrong by accident. It was pointed at the products that could carry it, and at aggregate level that is invisible, because the under-forecast on the constrained line and the over-forecast on the substitutes largely cancelled each other out.

The customer was on Vendor Managed Inventory, and that is the detail that turned a planning problem into a P&L one. Under VMI an overstated forecast does not sit in your own warehouse where somebody eventually walks past it and asks a question. It goes directly into the customer’s stock, on your paperwork, at your risk. It sat there for a year. At year end it came back as a large return of expired goods, and the SKUs in that return were exactly the ones the bias had been quietly loading up the whole time.

Nobody was hiding anything. Accuracy at the aggregate looked defensible from start to finish, the bias sign was pointing at under-forecasting, and the one cut that would have caught it, bias by product, was not on the report.

When is bias the right headline KPI?

Bias earns the top slot when the portfolio is simple enough that aggregate deviation is a fair proxy for how the supply chain actually behaved. Low product complexity, short or uniform lead times, delayed differentiation, a business where a customer order can be served from a near neighbour of the item that was forecast. Near-zero bias in that setting usually does mean efficiency.

It is also the right first KPI for organisations new to demand planning and S&OP. It tells them where they stand, in one number their leadership already understands, and it exposes whether the organisation runs on cautious pessimism or unrealistic optimism before anyone argues about statistical method.

What wrecks mean absolute error before the forecast even gets a chance

Mean absolute error is unforgiving, which is why so many businesses conclude that "the forecast is never accurate" and stop trying. Often the metric is measuring something other than forecast quality. Five reliable culprits:

What is really moving the numberHow it shows upWhy bias does not catch it
Master data qualityForecast and actuals attach to different item or location records, so error accumulates quietly between housekeeping cyclesIt usually nets off across periods, so the aggregate looks fine
Assumed lag vs demonstrated lead timeAccuracy is computed at a lag the supply chain does not actually run at, especially where lead times differ by productThe error cancels in the preceding or following period
Inventory coverHigh stock lets you serve orders you never forecast, which is good service and terrible measured accuracy. Static cover norms under a moving route to market make this worse every dayAggregate deviation is unaffected
Service-level agreementsOrders placed inside an agreed lead time are excluded from service measurement, so service and accuracy stop agreeing with each otherBias still measures the plan, not the exclusions
Portfolio typeFlexible portfolios (fertiliser, industrial chemicals, pack-size variants) serve one item from another in short orderFamily-level bias is the honest cut
Five things that move mean absolute error without telling you anything about forecast quality.

The rule underneath all five: if improving the KPI does not improve business delivery, the KPI is not a key performance indicator and it should be retired. In genuinely complex chains, packaged food and pharma being the obvious ones, mean absolute error is a life saver and belongs on the wall.

Inventory cover is the row we see doing the most damage, because it distorts mean absolute error every single day rather than once a quarter. Cover norms get set, and then reviewed on some annual cycle, while the route to market underneath them keeps moving. A norm that made sense for a two-week lane does not survive that lane changing. What you end up with is a shock absorber leaking oil. It still looks like a shock absorber and it holds up fine on smooth road. One real bump puts you in a ditch.

Cover has to be dynamic to work as a safety net. Any variability in supply or demand should move it, in either direction, and the good news matters as much as the bad: a lane that got faster or a supplier that got more reliable means you are sitting on cover you no longer need, serving orders you never forecast off the back of it, and wondering why the accuracy number looks the way it does. Left static, cover stops absorbing the noise and starts generating it.

Does volatility belong in the KPI set?

It does, and it is the only one of these numbers that needs no forecast to calculate. Coefficient of variation, the standard deviation of 90-day sales volume divided by its mean, tells you how erratic an item is and therefore how wrong its forecast is likely to be before you have forecast anything. Volatile items cost the chain real money in inventory, capacity or service, so measuring volatility and actively reducing it does more good than another round of model tuning.

The exception is a business model built on short-term opportunity. In commodity trading the whole point is to sell when the price is right, and unless you have econometric models good enough to forecast that, there is not much value in measuring forecast accuracy at all.

What do you do with a bias number once you have it?

Cut it two ways before you act on it, by product and by customer, because one aggregate number lets over-forecasting in one channel hide under-forecasting in another. Then look at who touches the forecast between the statistical output and the agreed plan, because persistent bias is almost always an incentive wearing a modeling costume. That handover is a governance question rather than a modelling one. We covered the fix in full in how to actually reduce forecast bias, and the metric choices around it in MAPE vs WMAPE vs WAPE.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *