Forecast Bias: The Formula, the Sign Trap, and When It Beats Chasing Accuracy
The forecast bias formula is the sum of actuals minus the sum of forecast, divided by the sum of forecast, over the same periods and the same lag: (ΣA − ΣF) ÷ ΣF × 100%. A positive result means you under-forecast and sold more than you planned, a negative result means you over-forecast, and anything that stays the same sign for several periods in a row is bias rather than noise. Bias matters more than the size of your error, because safety stock is designed to absorb random misses and is defenceless against directional ones.
What is the forecast bias formula?
Add up what actually happened, subtract the forecast, and divide by the forecast. That is it. The arithmetic is trivial and the interesting decisions all sit around it: which periods, which lag, which level of aggregation, and which sign convention. Bias measured at one lag on one aggregated total is a number that will comfort you. Bias measured at your operational lag, per product and per customer, is a number that will start arguments worth having.
The mean signed error is the same measurement without the percentage, (ΣA − ΣF) ÷ n, and it is often more useful when you are looking at a single item in units.
Which sign convention should you use?
Four are in common use and they disagree with each other. Getting two of them into the same meeting is how a demand review spends twenty minutes discovering that everybody agrees.
| Convention | Formula | A positive number means | Where we usually meet it |
|---|---|---|---|
| Deviation over forecast (our house standard) | (ΣA − ΣF) ÷ ΣF | We under-forecast, actuals beat the plan | Sales and finance reporting, S&OP attainment views |
| Error over actuals | (ΣF − ΣA) ÷ ΣA | We over-forecast | Demand planning packs, most planning software |
| Mean signed error | (ΣA − ΣF) ÷ n | We under-forecast, expressed in units | Item-level review, safety stock sizing |
| Forecast attainment index | ΣA ÷ ΣF | Above 1.0 means actuals beat the plan | Commercial reviews, supplier scorecards |
Pick one, write it at the top of the report, and never let the other three appear in the same deck without a label.
Ours is deviation over forecast: (ΣA − ΣF) ÷ ΣF, where a positive number means we sold more than we planned. Write that sentence on the report itself, in words, not just the symbols. The failure we run into most often is not a wrong formula. It is a page that prints a signed percentage and never says which way the sign points, so half the room reads it backwards and nobody finds out until somebody sizes a buffer on it.
We standardise this way because it is the number the commercial side of the business already speaks. A demand review that sits next to a sales review has to agree what "we were 6% out" means before it can agree on anything else, and deviation against plan is the version both rooms turn up already using. It is also the convention this site has run on since the 2016 piece this one follows from, so the back catalogue and the new work read the same way round.
One thing does move, and it is worth being straight about because it is our own back catalogue. Mean absolute error is often written as Σ|F − A| ÷ Forecast, and that is how the older pages here wrote it. Dividing by the forecast measures error against the plan you made, which flatters a business that plans low, punishes one that plans high, and moves the denominator every time the forecast moves. Our house standard divides by actuals, Σ|F − A| ÷ ΣA, as a percentage of what we actually sold. That gives you WAPE, it anchors the number to what really happened, and it stays comparable across periods and across items. Both versions are in live use across the industry, and you will meet the forecast-denominator one in plenty of places, our own earlier writing included (the forecast accuracy chapter of Hyndman and Athanasopoulos sets out the scale-free error measures this family sits in). We have standardised on actuals from here on and left the older pieces as they were published, rather than quietly rewriting the back catalogue to match a decision taken later.
Worked example: the same six months, two different answers
Worked example built to show the mechanics, not a client data set.
| Month | Forecast | Actual | Error (A − F) | Absolute error |
|---|---|---|---|---|
| Jan | 1,000 | 940 | −60 | 60 |
| Feb | 1,100 | 1,010 | −90 | 90 |
| Mar | 950 | 900 | −50 | 50 |
| Apr | 1,200 | 1,150 | −50 | 50 |
| May | 1,050 | 980 | −70 | 70 |
| Jun | 1,150 | 1,080 | −70 | 70 |
| Total | 6,450 | 6,060 | −390 | 390 |
House standard, deviation over forecast: −390 ÷ 6,450 = −6.0%. Negative, so we over-forecast in every month of the half and carried the stock for it. Run the same six months as error over actuals and the number is 390 ÷ 6,060 = +6.4%. Same spreadsheet, same half, opposite sign and a different magnitude. Neither is wrong. Both in one room with neither one labelled is how a business spends a quarter arguing about a direction it had already agreed on.
Is it bias or is it noise?
Compare the size of the signed total to the size of the absolute total. Call it the bias ratio: |Σ error| ÷ Σ |error|. It runs from 0 to 1. Near zero means your misses cancel out and you are looking at noise, which is what safety stock exists to cover. Near one means almost every unit of error points the same way and the buffer will never catch up.
In the table above, |−390| ÷ 390 = 1.00. Every month landed on the same side of the forecast. That is as pure as bias gets.
Now the same MAE with the error scattered:
| Month | Forecast | Actual | Error (A − F) | Absolute error |
|---|---|---|---|---|
| Jan | 1,000 | 1,060 | +60 | 60 |
| Feb | 1,100 | 1,010 | −90 | 90 |
| Mar | 950 | 1,000 | +50 | 50 |
| Apr | 1,200 | 1,150 | −50 | 50 |
| May | 1,050 | 1,120 | +70 | 70 |
| Jun | 1,150 | 1,080 | −70 | 70 |
| Total | 6,450 | 6,420 | −30 | 390 |
Identical mean absolute error of 65 units. Bias ratio |−30| ÷ 390 = 0.08. The first business needs a governance conversation. The second one needs a bigger buffer or a better model, and no amount of shouting at planners will move it. Same error size, completely different problem, and a single accuracy KPI cannot tell them apart.
This is the same idea as the classic tracking signal (cumulative signed error divided by mean absolute deviation), scaled to sit between 0 and 1 so it reads the same on any item. Use whichever your system already produces. Just make sure one of them is on the page.
What does low-grade forecast bias actually cost?
Enough to eat a year of margin on a product line, and it can do that while the headline bias number is pointing the other way.
We watched this happen at a business whose aggregate bias read as consistent under-forecasting. On the face of it that is the safer direction to be wrong in. It usually means you are leaving a bit of service on the table rather than money. What the aggregate was really reporting was one product line with a production capacity constraint, where genuine demand ran well ahead of anything anyone was willing to put in the forecast.
The sales team knew about the constraint. They also had a target. Since the SKU in real demand could not be made in the volume that target needed, the forecast went onto the SKUs that could. The number was not wrong by accident. It was pointed at the products that could carry it, and at aggregate level that is invisible, because the under-forecast on the constrained line and the over-forecast on the substitutes largely cancelled each other out.
The customer was on Vendor Managed Inventory, and that is the detail that turned a planning problem into a P&L one. Under VMI an overstated forecast does not sit in your own warehouse where somebody eventually walks past it and asks a question. It goes directly into the customer’s stock, on your paperwork, at your risk. It sat there for a year. At year end it came back as a large return of expired goods, and the SKUs in that return were exactly the ones the bias had been quietly loading up the whole time.
Nobody was hiding anything. Accuracy at the aggregate looked defensible from start to finish, the bias sign was pointing at under-forecasting, and the one cut that would have caught it, bias by product, was not on the report.
When is bias the right headline KPI?
Bias earns the top slot when the portfolio is simple enough that aggregate deviation is a fair proxy for how the supply chain actually behaved. Low product complexity, short or uniform lead times, delayed differentiation, a business where a customer order can be served from a near neighbour of the item that was forecast. Near-zero bias in that setting usually does mean efficiency.
It is also the right first KPI for organisations new to demand planning and S&OP. It tells them where they stand, in one number their leadership already understands, and it exposes whether the organisation runs on cautious pessimism or unrealistic optimism before anyone argues about statistical method.
What wrecks mean absolute error before the forecast even gets a chance
Mean absolute error is unforgiving, which is why so many businesses conclude that "the forecast is never accurate" and stop trying. Often the metric is measuring something other than forecast quality. Five reliable culprits:
| What is really moving the number | How it shows up | Why bias does not catch it |
|---|---|---|
| Master data quality | Forecast and actuals attach to different item or location records, so error accumulates quietly between housekeeping cycles | It usually nets off across periods, so the aggregate looks fine |
| Assumed lag vs demonstrated lead time | Accuracy is computed at a lag the supply chain does not actually run at, especially where lead times differ by product | The error cancels in the preceding or following period |
| Inventory cover | High stock lets you serve orders you never forecast, which is good service and terrible measured accuracy. Static cover norms under a moving route to market make this worse every day | Aggregate deviation is unaffected |
| Service-level agreements | Orders placed inside an agreed lead time are excluded from service measurement, so service and accuracy stop agreeing with each other | Bias still measures the plan, not the exclusions |
| Portfolio type | Flexible portfolios (fertiliser, industrial chemicals, pack-size variants) serve one item from another in short order | Family-level bias is the honest cut |
The rule underneath all five: if improving the KPI does not improve business delivery, the KPI is not a key performance indicator and it should be retired. In genuinely complex chains, packaged food and pharma being the obvious ones, mean absolute error is a life saver and belongs on the wall.
Inventory cover is the row we see doing the most damage, because it distorts mean absolute error every single day rather than once a quarter. Cover norms get set, and then reviewed on some annual cycle, while the route to market underneath them keeps moving. A norm that made sense for a two-week lane does not survive that lane changing. What you end up with is a shock absorber leaking oil. It still looks like a shock absorber and it holds up fine on smooth road. One real bump puts you in a ditch.
Cover has to be dynamic to work as a safety net. Any variability in supply or demand should move it, in either direction, and the good news matters as much as the bad: a lane that got faster or a supplier that got more reliable means you are sitting on cover you no longer need, serving orders you never forecast off the back of it, and wondering why the accuracy number looks the way it does. Left static, cover stops absorbing the noise and starts generating it.
Does volatility belong in the KPI set?
It does, and it is the only one of these numbers that needs no forecast to calculate. Coefficient of variation, the standard deviation of 90-day sales volume divided by its mean, tells you how erratic an item is and therefore how wrong its forecast is likely to be before you have forecast anything. Volatile items cost the chain real money in inventory, capacity or service, so measuring volatility and actively reducing it does more good than another round of model tuning.
The exception is a business model built on short-term opportunity. In commodity trading the whole point is to sell when the price is right, and unless you have econometric models good enough to forecast that, there is not much value in measuring forecast accuracy at all.
What do you do with a bias number once you have it?
Cut it two ways before you act on it, by product and by customer, because one aggregate number lets over-forecasting in one channel hide under-forecasting in another. Then look at who touches the forecast between the statistical output and the agreed plan, because persistent bias is almost always an incentive wearing a modeling costume. That handover is a governance question rather than a modelling one. We covered the fix in full in how to actually reduce forecast bias, and the metric choices around it in MAPE vs WMAPE vs WAPE.
