• Research Paper
  • Current
  • Advanced
  • 15 min read

What a Forecast Can and Cannot Do

A forecast is not a prediction of demand. It is an input to an ordering decision, and the difference explains how a company can get its revenue number almost exactly right and still lose 87% of its operating income.

Quantitative Methods · Retail

Key takeaways

  • About 50% of individual business forecasts fail to beat a plain “same as last period” naive forecast (Morlidge, Foresight 33, 2014). Until you run that benchmark you do not know whether yours adds anything at all.
  • Across the M4 competition's 100,000 series, 12 of the 17 most accurate methods were combinations, and none of the six pure machine-learning entries beat the combination benchmark; only one beat Naïve2.
  • Target's second quarter of fiscal 2022 is what a right forecast and a wrong plan look like: total revenue rose 3.5% year over year to $26.04B while operating income fell 87% to $321M, clearing inventory that had run 43% above the prior year.
  • Safety stock is priced in z-values, not service levels. Illustrative only: buying 90%→95% cycle service costs about $79 per point of service; 95%→99% costs about $185 per point; and cutting lead time from four weeks to one delivers 99% service for less than 90% used to cost.

A forecast is an input to a decision, not a prediction

The useful question is never “what will demand be.” It is “what should I order, staff, or build, given that I do not know what demand will be.” A forecast that improves that decision has done its job even when the number turns out wrong; a forecast that changes no decision has done nothing, however close it lands.

This reframing has a hard consequence: a forecast can only be judged against an alternative. The alternative is not a competing model, it is the cheapest thing you could have done instead: carry last period's number forward. Steve Morlidge, working across supply-chain forecast data and writing in Foresight in 2014, put the finding bluntly: many companies fail to achieve a relative error better than a simple “same as last period” naive forecast, and around 50% of individual forecasts fail to meet that benchmark. Half. Not half of companies, but half of the individual line-item forecasts, produced by people paid to produce them, inside processes with meetings and sign-offs attached.

The M competitions say the same thing from the research side. M4 put 61 methods against 100,000 real series. Its published findings are worth reading exactly: 12 of the 17 most accurate methods were combinations of mostly statistical approaches; the winner, a hybrid using both statistical and machine-learning features, was close to 10% more accurate than the combination benchmark on average sMAPE; and the six pure machine-learning methods performed poorly, with none more accurate than the combination benchmark and only one more accurate than Naïve2.

Read that last clause slowly. Six serious machine-learning submissions, and five of them lost to a benchmark that barely qualifies as a method. The lesson is not that machine learning does not work, because the winner used it. The lesson is that sophistication is not the axis that matters, and that averaging several unglamorous methods beats picking one clever one. If you take a single practical rule from the forecasting literature, take that one: combine, and always score against naive.

So the honest boundary. A forecast can tell you the central tendency of a stable process, the shape of a seasonal pattern you have seen several times, and, most valuably, how wide the uncertainty is. It cannot tell you when a regime breaks, because the break has by definition not appeared in the history the model was fitted on. Every forecast is a bet that next period resembles the periods it learned from.

The error metric you choose is the behaviour you will get

Pick a scoring rule and you have picked a bias. This is not a subtlety; it is the single most common way a competent forecasting team quietly produces a bad plan.

The common measures split into three families. Scale-dependent errors (mean absolute error, root mean squared error) are in the units of the thing you sell, which makes them readable but useless for comparing a fast SKU against a slow one. They also disagree about what a forecast is: minimising MAE gives you the median outcome, minimising RMSE gives you the mean. If your demand distribution is skewed, and retail demand almost always is, those are different numbers, and you will hold different inventory depending on which you optimised.

Percentage errors solve comparability and create worse problems. Hyndman and Athanasopoulos, in Forecasting: Principles and Practice, list the defects of MAPE precisely: it is infinite or undefined when the actual is zero, it produces extreme values when the actual is near zero, it assumes the measurement scale has a meaningful zero, and it puts a heavier penalty on negative errors than on positive ones. That last one is the operational landmine. With MAPE defined against the actual, the worst possible under-forecast is bounded, since forecasting zero against an actual of 100 scores 100%, while over-forecasting is unbounded: forecast 300 against an actual of 100 and you score 200%. A team scored on MAPE learns, without anyone deciding it, to forecast low. Forecasting low causes stockouts. The metric that was supposed to make the plan accurate has made it systematically short.

The symmetric variant does not rescue it. sMAPE, defined as the mean of 200 × |y − ŷ| / (y + ŷ), still divides by near-zero quantities and can take negative values, which is an odd property for an absolute-error measure; the same authors recommend it not be used. Their recommendation is the mean absolute scaled error, which divides each error by the average training-set error of a naive forecast. MASE is scale-free, comparable across series, and, in the part that matters, has the naive benchmark built into its denominator. A MASE above 1 means you did worse than doing nothing. You do not have to remember to run the comparison; the number is the comparison.

The other half of the job is separating bias from spread. Mean error, keeping the sign, tells you whether you are consistently over or under. Mean absolute error tells you how far off you are. They are independent problems with independent fixes: bias is a process defect (a sales team sandbagging, a planner adding a private buffer, a MAPE target pushing everyone low) and it is nearly free to remove once measured. Spread is irreducible noise and gets handled by inventory, not by modelling. Teams that track only absolute error spend years attacking the expensive half of the problem while the cheap half sits uncorrected in plain sight.

Right at the top, wrong at every line

Forecast error does not average out the way intuition says it does, and the direction of the mistake is the opposite of what most planners assume.

Start with the arithmetic that makes aggregate forecasts look good. Illustrative only: a distributor carries 20 SKUs, each averaging 400 units a week with a standard deviation of 120, a coefficient of variation of 30% on every single line. Add them up. If the SKUs move independently, the total averages 8,000 units a week with a standard deviation of 120 × √20 = 537, a coefficient of variation of 6.7%. The mean scaled by 20; the noise scaled by 4.47. The company forecast is now four and a half times more stable than any of the decisions inside it.

This is why the monthly business review feels calm while the warehouse is on fire. Nobody orders inventory at the aggregate. Every purchase order is placed against a single SKU with a 30% coefficient of variation, and the total being right is no comfort to the buyer who ordered 2,000 of the wrong colour.

Target's fiscal 2022 is the public version of exactly this. Its 10-Q filings put merchandise inventory at $15.08B at 30 April 2022 against $10.54B a year earlier, 43% higher. What happened next was not a revenue miss. For the quarter ended 30 July 2022 the company reported total revenue of $26.04B, up 3.5% from $25.16B a year before. Demand, in aggregate, was almost exactly where it should have been. Operating income for that same quarter was $321M, against $2.47B the year before: down 87%, an operating margin of 1.2% where the prior year had produced 9.8%. Inventory did not stop climbing until it peaked at $17.12B in the October quarter, and did not normalise to $13.50B until the fiscal year closed in January 2023.

A revenue forecast within a few percent, and a plan that cost roughly $2.1B of operating income in a single quarter. The error was never in the total. It was in mix, meaning which categories, which sizes and which weeks, and mix is precisely what aggregation hides. The practical rule: forecast at the level you decide at, and if you must forecast at a higher level, reconcile back down explicitly rather than allocating by last year's share, because last year's share is itself a forecast.

Safety stock is where the forecast error gets priced

You cannot forecast your way out of uncertainty; you can only decide how much of it to buy insurance against. Safety stock is that premium, and it has a formula that makes the price legible.

The standard construction, set out in Peter King's APICS treatment of the subject, is safety stock = Z × σ over the replenishment period, where Z is the z-value of the target cycle service level. When demand variability dominates and the performance cycle differs from the interval you measured σ over, that becomes Z × √(PC/T₁) × σ_D. When lead-time variability dominates instead, it is Z × σ_LT × D_avg. Where both are present and independent, they combine under the root of the sum of squares, which is always less than adding them, a small mercy worth taking.

Work it. Illustrative only: one SKU, weekly demand averaging 400 units with a weekly standard deviation of 120, a fixed four-week supplier lead time, a $18 unit cost and a 25% annual carrying charge, so $4.50 per unit per year. Demand over the lead time averages 1,600 units; its standard deviation is 120 × √4 = 240.

At 90% cycle service, z = 1.28, so safety stock is 307 units and the reorder point is 1,907. At 95%, z = 1.645, safety stock is 395 and the reorder point is 1,995. At 99%, z = 2.33, safety stock is 559 and the reorder point is 2,159. Carrying costs: $1,382, $1,778 and $2,516 a year respectively. Moving from 90% to 95% costs $396 a year, about $79 per point of service, and removes half the stockout cycles. Moving from 95% to 99% costs a further $738, about $185 per point, more than twice the price for a smaller and smaller sliver of protection. This is the shape of the whole curve: the tail is expensive and gets worse, because z climbs without limit as the service level approaches 100.

Now the lever almost nobody pulls. Keep everything else and cut the lead time from four weeks to one. The standard deviation over the lead time falls from 240 to 120, and 99% service now needs 2.33 × 120 = 280 units of safety stock, costing $1,260 a year, which is less inventory and less money than 90% service used to require at a four-week lead time. Halving the square root of lead time beats any amount of statistical tuning. Before you argue about the model, get a quote on a faster shipping mode or a nearer supplier and compare it against the carrying cost you would avoid.

One warning that trips up almost every first implementation: cycle service level is not fill rate. Cycle service level is the fraction of replenishment cycles that end without a stockout: it counts how often, and says nothing about how badly. Fill rate is the fraction of demanded units actually supplied. A SKU can run a 95% cycle service level and a wretched fill rate if the stockouts that do happen are deep and long. Customers experience fill rate. Report the one they feel.

You are forecasting sales, not demand — and the chain amplifies it

The data in your system is not what customers wanted. It is what you managed to sell them. Every stockout censors the record: demand that arrived while the shelf was empty shows up as a zero, or as a smaller number, or as nothing at all.

The feedback loop that follows is nasty and quiet. The model fits the censored series, so the estimated mean drifts down. A lower mean sets a lower reorder point. A lower reorder point produces more stockouts. More stockouts censor more demand. Nothing in the pipeline is broken and no alert fires; the SKU simply starves, and the report shows a well-behaved forecast with a small error, because the model is now predicting your own supply constraint with great precision. If you run lost-sales substitution, flagging stockout periods and imputing rather than fitting the zero, do it before you touch anything else in the model. It is usually the largest single accuracy gain available.

Upstream, the same distortion compounds. Lee, Padmanabhan and Whang gave the effect its name and its mechanism in Management Science in 1997: Procter & Gamble found that diaper orders from distributors varied more than consumer demand could explain, and Hewlett-Packard found that reseller orders to its printer division swung far harder than end-customer demand. Their result is that variance amplification arises when each tier updates its order-up-to level from the demand signal it observes, and the amplification grows with replenishment lead time, and, notably, exists even when the lead time is zero. Rationing games, order batching and price promotions each add their own layer on top.

The macro data shows the whole system doing this at once. The Census Bureau's manufacturing and trade series has total business inventories-to-sales sitting at 1.43 in February 2020, spiking to 1.74 in April 2020 as sales vanished faster than goods could stop arriving, then collapsing to 1.26 by April 2021 during the shortage, then climbing all the way back to 1.43 by December 2022 during the glut. Retail was sharper still: 1.43, then 1.68, then a low of 1.09 in June 2021, then 1.28 by December 2022. As of June 2026 the totals sit at 1.30 and 1.25.

Three full swings in under three years, in a series that normally moves by hundredths. No forecasting method available to any firm in 2020 was going to call that. What separated the firms that survived it was not prediction. It was how quickly they could change an order after the signal arrived.

Planning for a forecast you know will be wrong

Once you accept that the number will be wrong, the design problem changes from “make it more accurate” to “make being wrong cheaper.” That is a much more tractable engineering problem, and it has a short list of levers.

Shorten the commitment horizon. Every week of lead time is a week of demand you had to guess at, and the safety-stock arithmetic above prices it directly. Split orders: commit half at the long lead time and half at a shorter, dearer one, and let the second tranche be informed by three more weeks of actual sales. The premium on the fast half is usually far less than the markdown on the units you would otherwise have guessed wrong.

Move the decoupling point. King's own list of alternatives to holding more safety stock is instructive: expedite selectively rather than blanket-buffer, and consider make-to-order or finish-to-order production where lead times allow. Finish-to-order in particular works because it lets you hold the generic thing, which has a low coefficient of variation because it pools across every variant, and defer the specific thing until the order exists. This is aggregation used deliberately, in inventory rather than in the spreadsheet, and it is the only place aggregation genuinely helps.

Hold options instead of goods. A supplier agreement with a firm minimum and a priced option on additional units converts a forecasting problem into a pricing problem, which is easier. So does a return or markdown-money clause, which is the same thing with the risk sitting one tier up.

And instrument the review. Every forecast cycle should produce four numbers, not one: the naive benchmark's error, your error, the bias with its sign, and the number of SKUs where you lost to naive. The last one is the diagnostic that actually changes behaviour, because it is a list of names rather than an average, and the correct response to a SKU you cannot beat naive on is to stop forecasting it and use naive, which frees the planner's attention for a SKU where judgement is worth something.

The deepest trap is the one that closes the loop with the incentive: a forecast is often owned by someone whose bonus depends on hitting it. That person will not give you their best estimate. They will give you a number they are confident of exceeding, and they will be right about it, and every downstream inventory decision will be built on a deliberate understatement. If your forecast accuracy has been suspiciously stable and suspiciously good for a long time, check what happens to the forecaster when the number is missed before you check the model.

Put it to work

Run the naive benchmark first: score last quarter's forecasts against “same as last period” and keep only the SKUs where you beat it. Then split error into bias and spread, and track bias separately. Recompute safety stock from the standard deviation of demand over lead time, not a weeks-of-cover rule. Ask the supplier what a shorter lead time costs before buying more service level.

Sources & references

Linked entries open the named source directly. Entries without a link say exactly what kind of reference they are — and how to check them yourself.

Educational note: This briefing is general business education, not financial, legal, tax, or investment advice. Figures and rules change and vary by situation — verify current specifics with primary sources and qualified professionals before acting.