MAPE vs WMAPE vs WAPE: When Each Forecast Metric Lies to You
If you only track one forecast accuracy metric, make it WAPE (weighted absolute percentage error, sometimes sold as WMAPE): total absolute error divided by total actual demand. It survives mixed-volume portfolios and near-zero actuals, which are the two situations where plain MAPE quietly falls apart. Every one of these metrics lies under specific conditions, though, and knowing when matters more than picking the right one.
What’s the difference between MAPE, WMAPE and WAPE?
MAPE averages each period or SKU’s percentage error with equal weight, so a C-item that missed by 200% counts the same as your top seller missing by 5%. WAPE is the sum of absolute errors divided by the sum of actuals, which makes it volume-weighted by construction. WMAPE is WAPE with explicit weights, usually volume or revenue. In most software the last two are the same calculation, and the naming difference is vendor marketing more than mathematics.
When does MAPE lie?
Three reliable failure modes. First, near-zero actuals: one SKU that sold 1 unit against a forecast of 5 posts a 400% error and wrecks the average, and when actuals are zero the formula divides by zero. Second, the long tail: average enough C-item noise into MAPE and the number says your forecasting is terrible while the warehouse runs fine. Third, asymmetry: MAPE punishes over-forecasting harder than under-forecasting, so a team managed on MAPE learns to forecast low. You then pay for that lesson in stockouts.
Run one small portfolio through both metrics and the problem is obvious:
| SKU | Forecast | Actual | Abs error | APE |
|---|---|---|---|---|
| A-item | 1,000 | 950 | 50 | 5.3% |
| C-item 1 | 20 | 5 | 15 | 300% |
| C-item 2 | 30 | 20 | 10 | 50% |
| C-item 3 | 15 | 10 | 5 | 50% |
| B-item | 25 | 35 | 10 | 28.6% |
| MAPE (average of APEs) | 86.8% | |||
| WAPE (90 ÷ 1,020) | 8.8% |
Same five SKUs, same month. One metric says the process is broken, the other says the warehouse barely noticed. Both are telling a partial truth, which is the point.
The bars on the left are sized by percentage error. C-item 1 missed by 300%, so it towers over the A-item’s quiet 5.3% miss, and MAPE treats both misses as equally important. The bars on the right are sized by each SKU’s share of total volume. The A-item carries 93.1% of the portfolio’s actual demand, so its small error is what actually moves WAPE, while the C-items’ dramatic percentage misses barely register. Same five orders. Two honest but incomplete stories.
When does WAPE lie?
WAPE’s weighting is also its blind spot. Your A-items dominate the number, so WAPE can read 9% while the tail, where the write-offs and stockouts actually live, is unforecastable chaos. A good WAPE sitting on top of a rotten tail is a common and expensive combination. Segment before you average. WAPE by ABC class tells you something. One company-wide WAPE mostly hides things.
Why isn’t bias covered by MAPE or WAPE?
Accuracy metrics measure the size of the error. None of them tell you its direction, and a forecast that runs persistently 10% high with a beautiful WAPE will still fill your warehouse. Track bias alongside whichever accuracy metric you choose. The pair costs nothing extra and each catches what the other misses.
Why do teams struggle with these KPIs at all?
In all honesty, corporations struggle with forecast accuracy in general. Forecasting is a mathematical process and the KPI guardrails around it are statistical measures, which leaves non-technical functions with three hurdles: understanding the KPI, visualizing its practical implication, and devising ways to improve it. The funnel of interest dwindles most at the second hurdle. If the cross-functional team that owns the demand and supply levers cannot see what forecast accuracy does to their own outcomes, they will certainly not act to improve it. Picture a salesperson carrying a stretch target with commission riding on it. Forecast accuracy is the last thing on their mind, and no dashboard changes that.
How do you make these KPIs actionable?
Cut one: lags that mirror your supply chain. Lags are the easiest part of this toolkit to understand, and here is how we use them. Always keep several in play. If it takes four weeks from receiving an order to getting stock on a customer’s shelf, then accuracy at a four-week lag is a good estimate of why service is good or bad. If the total supply chain reaction time is thirteen weeks, from raw material ordering to stock on the customer’s shelf, the thirteen-week lag shows how responsive the chain really is. And if capacity takes a year to ramp, the 52-week lag tells you how well capex management can trust the long-range demand signal. Each lag answers a different business question, which is exactly why one lag is never enough.
Cut two: predictability classes. ABC classification weights accuracy by portfolio contribution; confidence intervals do the same for inherent predictability. A product with erratic, impulsive demand (think chocolate bars at the till) earns a wide interval and a lower accuracy expectation. Laundry detergent or airline seats follow far more predictable patterns and should be forecastable with high confidence. Cut the KPI this way and misses on predictable products stop hiding behind the chaos of impulsive ones. Those misses are usually systemic process issues, and systemic issues are normally fixable.
So what should you actually track?
WAPE plus bias, segmented by ABC class and predictability, measured at the lags your decisions use, meaning the forecast you bought against rather than the one somebody fixed last week. In our experience that lag discipline moves more organizations from arguing to improving than any choice of metric ever has. See our piece on why S&OP fails for what happens when nobody agrees on which number the room is even arguing about.
