AI Demand Forecasting: From Spreadsheet Guesswork to Production-Grade Predictive Planning

What Is AI Demand Forecasting?
AI demand forecasting uses machine learning models trained on historical demand plus external signals to predict future demand. It replaces spreadsheet averages and manual adjustment.
Two things people mash together, worth separating up front:
• Forecasting produces a number.
• Demand planning is what the business decides to do with that number, how much to buy, build, or staff for.
AI here isn’t one thing. It’s three families of method, and picking between them is most of the actual skill. That’s really what this article is about.
Why Spreadsheet Forecasting Breaks Down
Spreadsheets are genuinely fine for a handful of stable products. They stop working at a specific point, and that point shows up fast once any of these apply:
• Hundreds or thousands of SKUs
• Promotions and price changes that shift demand unpredictably
• Items that sell zero most weeks
• New products with no history to average
• Nobody can hand-tune four thousand series every month without losing their mind
You know the symptom already if you’ve lived it: the forecast that gets “adjusted” by three different people before anyone actually acts on it.
| Spreadsheet approach | What breaks | What a model does instead |
| One average or trend line per SKU | Falls apart past a few hundred SKUs, let alone a few thousand | Learns one model across all series at once |
| Manual adjustment for promotions | Doesn’t scale, and depends on whoever remembers to do it | Promo flag as a feature, learned from past uplift |
| Flat zero-fill for slow-moving items | Treats real intermittent demand as noise or as nothing | Purpose-built methods (Croston’s) for lumpy demand |
| New product gets last year’s number, copied down | There is no last year for a new SKU | Forecast by analogy, or a zero-shot foundation model |
| “Adjusted” by three people before anyone acts on it | Nobody can say whether the adjustments helped or hurt | A tracked baseline everything else has to beat |
What Data You Actually Need
The Internal Data That Does Most of the Work
Most of what a model needs is already sitting in systems you already run. The floor:
• Transaction or POS history
• Order history from the ERP
• The price and promotion calendar
• Inventory and stockout records
• Product hierarchy
• A calendar of holidays and events
Roughly two to three years of history is the rough floor, enough to see seasonality repeat itself at least twice.
The one almost nobody puts in writing: stockout records matter more than people think. Demand you couldn’t serve is invisible in your sales data, all it shows is what got sold, not what people wanted.
A model trained on that censored history quietly learns to under-forecast, forever, and nobody notices until stock keeps running short even though “the forecast looks fine.” This is probably the most common silent failure in the whole field, and it costs nothing to fix if you catch it early.
The External Signals Worth Adding, and When They Aren’t
A handful of external signals genuinely help, but only in the right context:
• Weather
• Macroeconomic indicators
• Competitor pricing
• Search trends
• Local events
Weather-sensitive categories and short-horizon forecasts benefit the most. Stable industrial demand, the kind that doesn’t care what the search trend for umbrellas looks like, often gains nothing from any of it.
Every signal you add is also another pipeline that can quietly break at 2am. Add them because they earn their place, not because the vendor deck had a nice slide about it.
Getting the Data Into Shape
Before any model sees the data, three decisions need to be made deliberately rather than by default:
• Pick the aggregation level the decision is actually made at, daily or weekly, not whatever’s finest available just because it exists.
• Handle missing periods and outlier spikes like the one-off bulk order that isn’t going to repeat.
• Decide the hierarchy question early. Forecast at SKU level and roll up, or forecast the total and split down. This is hierarchical reconciliation: the set of methods that keeps your SKU-level and category-level forecasts from quietly contradicting each other.
The Three Families of Forecasting Method
This is the part worth spending real time on. If you read one section of this article carefully, make it this one.
Classical Statistical Methods: ARIMA, ETS, Croston
Three methods, three different jobs, and all three remain genuinely competitive:
• ETS earns its keep on short, clean, seasonal series where there isn’t much noise to model.
• ARIMA is the right call when the autocorrelation structure actually matters, when what happened a few periods ago genuinely predicts what happens next.
• Croston’s method exists specifically for intermittent demand, the huge share of real spare-parts and B2B catalogues that sell in fits and starts, and almost no AI-forecasting article bothers to mention it.
They’re fast, cheap to run, and you can explain exactly why they said what they said, which is worth more in a planning meeting than people give it credit for.
Machine Learning: Gradient Boosting and Neural Approaches
Two tiers here, and most teams only need the first one:
• LightGBM and XGBoost, fitted on engineered features (lags, rolling means, calendar signals, price, promo flags), are the workhorse.
• LSTM, DeepAR and the Temporal Fusion Transformer are the deep-learning options for when you have enough data and enough patience to tune them properly.
Here’s the point most of this SERP misses entirely: ML wins mainly through cross-learning. One model trained across thousands of related series lets a sparse, low-volume item borrow statistical strength from similar ones that sell more often.
That’s the actual mechanism. It’s not that gradient boosting is smarter than ETS in some abstract sense, it’s that it can see patterns across a whole catalogue that a series-by-series model never gets to look at.
This isn’t a theoretical claim. In the M5 forecasting competition, run on real Walmart retail data, LightGBM was used by all of the top 50 competitors, though the same results also found simple exponential smoothing still competitive at the individual product level, which is exactly the nuance the cheat sheet below tries to capture.
Here’s a working version, feature engineering plus a proper walk-forward split rather than a random one, since a random split lets future data leak into training and quietly flatters the result:
# pandas 2.2, lightgbm 4.7.0 (tested Sep 2026)
import pandas as pd, lightgbm as lgb
# sales: columns = sku, week, demand
sales = sales.sort_values(["sku", "week"])
sales["lag_1"] = sales.groupby("sku")["demand"].shift(1)
sales["lag_52"] = sales.groupby("sku")["demand"].shift(52)
sales["roll_mean_4"] = sales.groupby("sku")["demand"].shift(1).rolling(4).mean()
sales["week_of_year"] = sales["week"].dt.isocalendar().week.astype(int)
data = sales.dropna(subset=["lag_1", "roll_mean_4"])
cutoff = data["week"].quantile(0.8) # walk-forward, not random
train, test = data[data.week <= cutoff], data[data.week > cutoff]
features = ["lag_1", "roll_mean_4", "week_of_year"]
model = lgb.LGBMRegressor(n_estimators=200, learning_rate=0.05)
model.fit(train[features], train["demand"])
preds = model.predict(test[features])
wape = (test["demand"] - preds).abs().sum() / test["demand"].sum() * 100
print(f"Walk-forward WAPE: {wape:.1f}%")
LightGBM documentation, official docs for the library used in the code block above.
Run against a synthetic 20-SKU weekly dataset with mild seasonality, that came out to a walk-forward WAPE a little under 5%, roughly the ballpark you’d hope for on well-behaved series with only three basic features.
Real retail data with real promotions will be messier, but the shape of the code doesn’t change much.
Time-Series Foundation Models: The 2026 Question
Pre-trained models, the Chronos, TimesFM, Moirai and Lag-Llama family, can forecast a brand-new series zero-shot, with no training on your data at all.
The appeal is real: near-zero setup, genuinely useful for cold-start items and as a fast first baseline before you’ve built anything proper.
The honest caveat, sourced rather than asserted: benchmarking these foundation models is a genuinely contested question, and how you evaluate them fairly is still being actively argued over in the literature.
Treat a foundation model as a strong starting point, not as a settled verdict on what will work best for your catalogue.
Which One Should You Actually Use?
Save this table. It’s the closest thing this article has to a cheat sheet.
| Situation | Starting point | Why, and what it must beat |
| A few stable, well-behaved series | ETS or a well-tuned ARIMA | Fast, cheap, explainable, and genuinely competitive here. Must beat: naive (last period’s actual). |
| Thousands of related SKUs | Gradient boosting (LightGBM/XGBoost) trained across all series | Cross-learning lets sparse items borrow strength from similar ones. Must beat: a per-SKU ETS model. |
| Intermittent, lumpy demand | Croston’s method or a variant (SBA) | Built for demand that’s zero most weeks. Must beat: flat average or zero-fill. |
| Brand-new product, no history | Forecast by analogy, or a foundation model zero-shot | There’s nothing to train on yet. Must beat: copying last year’s number for a similar item. |
| Strong promotional effects | ML with explicit promo/price features | Statistical methods don’t have a slot for this signal. Must beat: a model with no promo flag at all. |
| Very short-horizon operational forecasting | Simple statistical method, refreshed often | Complexity buys little at a one- or two-day horizon. Must beat: yesterday’s number, carried forward. |
Forecasting: Principles and Practice, Hyndman & Athanasopoulos, the open textbook behind the ARIMA, ETS and reconciliation references above.
How Forecast Accuracy Is Measured
MAPE, WAPE and MAE, and Why MAPE Lies
Here’s the arithmetic, worked in the open, on six periods of a single series:
| Period | Actual | Forecast |
| 1 | 120 | 110 |
| 2 | 95 | 100 |
| 3 | 140 | 150 |
| 4 | 10 | 20 |
| 5 | 0 | 5 |
| 6 | 130 | 125 |
| Period | Absolute error | % error |
| 1 | 10 | 8.3% |
| 2 | 5 | 5.3% |
| 3 | 10 | 7.1% |
| 4 | 10 | 100.0% |
| 5 | 5 | undefined |
| 6 | 5 | 3.8% |
MAPE, averaging the percentage errors across the five periods where actual demand wasn’t zero, comes out to 24.9%. WAPE, weighting each error by volume instead of averaging percentages blind, comes out to 9.1%.
Same data, wildly different story. Period 4 is genuinely off by 10 units against an actual of only 10, a 100% error that MAPE treats as equally important as period 1’s much larger absolute miss on a much bigger number.
MAPE is also flatly undefined in period 5, where actual demand was zero, you can’t divide by it. That’s exactly the kind of period retail and spare-parts data is full of, which is the whole reason planning teams lean on WAPE instead.
MAE here is 7.5 units, the plain average absolute miss, useful as a sanity check but not something you’d report to a room full of people comparing categories with wildly different volumes.
Bias: the Error Everyone Ignores
Bias is direction, not size. On this same table, the average signed error (forecast minus actual) is +2.5 units, meaning the forecast is running about 3% high overall.
A forecast can post a perfectly respectable MAPE or WAPE and still be consistently biased high, which quietly builds inventory all year without anyone noticing, until someone asks why the warehouse is fuller than it should be.
Bias is often the more expensive error of the two, precisely because it hides behind a decent-looking accuracy number.
Forecast Value Added: Did the Model Earn Its Place?
FVA compares your model against a naive baseline, last period’s actual, or the same period last year, whichever makes sense for the series.
The uncomfortable thing worth saying plainly: a lot of forecasting effort, including manual overrides stacked on top of a perfectly good model, produces negative FVA. It made things worse, confidently.
The rule that follows is simple and worth pinning above someone’s desk: if it doesn’t beat naive, it doesn’t ship.
Backtesting Properly
Two rules, and skipping either one quietly inflates every accuracy number you report:
• Walk-forward, or rolling-origin, validation only, never a random train-test split. A random split lets the model peek at the future while training on the past, and it only gets caught once it’s live.
• Match the test horizon to the actual decision horizon. Backtesting a one-week-ahead model on a 90-day horizon tells you very little about how it’ll behave when it matters.
Why AI Demand Forecasting Projects Fail
This is the section vendors are structurally unable to write, so here it is, straight, with a fix attached to each one.
• Censored demand baked into training data. Stockouts get read as low demand instead of unmet demand. Fix: bring in stockout records and treat those periods as missing, not zero, before you train anything.
• Forecasting at the wrong granularity for the decision. A daily forecast feeding a weekly ordering decision is solving a problem nobody has. Fix: match the aggregation level to where the decision actually gets made.
• No baseline, so nobody can tell if the model helped. Without FVA against naive, “the model works” is just an opinion. Fix: compute the naive baseline before you fit anything fancier, not after.
• Planners overriding every number because the model was never explained. Trust doesn’t show up on its own. Fix: show planners the reasoning, the features, the recent accuracy, not just a number that appeared from nowhere.
• A notebook model with no retraining pipeline. Great backtest, then six months of silent drift as the world moves and the model doesn’t. Fix: schedule retraining and monitor accuracy in production, not just at launch.
• One method applied across a whole catalogue, intermittent items included. Croston’s exists for a reason; ETS on a lumpy spare-parts series produces confident nonsense. Fix: segment the catalogue and route each segment to the method built for it.
• A point forecast treated as certainty. The business needs a range to plan safety stock against, not a single number presented as fact. Fix: carry a prediction interval, not just a point estimate, all the way through to the planning decision.
How to Run a Forecasting Project End to End
1. Define the decision the forecast actually serves, ordering, staffing, production scheduling, whatever it is
2. Set the horizon and the granularity to match that decision, not to whatever’s easiest to pull
3. Assemble and audit the data, stockouts and all, before touching a model
4. Build the naive baseline first, always, before anything more sophisticated
5. Fit a simple statistical model per series as a second baseline
6. Fit an ML model across the full set of related series
7. Backtest walk-forward, never on a random split
8. Measure Forecast Value Added against the naive baseline, honestly
9. Ship with monitoring and a retraining schedule already built in, not bolted on later
Notice where the baseline sits: before the model, not as an afterthought once someone asks whether it worked. Most teams do this backwards, and then genuinely cannot answer the one question that matters.
Skills and Roles Behind a Working Forecast
The skills, roughly in the order they actually matter on the job:
• SQL and Python
• Time-series thinking
• Feature engineering
• An honest grasp of validation, specifically why a random split lies to you
• Enough business context to know what the forecast is actually for
The roles this sits under, in practice: demand planner, supply chain analyst, data scientist, analytics manager, often with the work split across two or three of them rather than owned by one.
Everything above, the methods, the metrics, the failure modes, only matters once someone can actually run it end to end and defend the result in a planning meeting. If that’s the gap you’re closing next, AI-Powered Operations is a module inside Scaler’s PGP in Business & AI, built around exactly this: forecasting, planning, and the judgement calls that sit between a model’s output and a business decision.
Frequently Asked Questions
Using machine learning models trained on historical demand plus external signals to predict future demand, instead of spreadsheet averages and manual adjustment. See the opening definition for the fuller version.
Often, but not always. It wins most clearly across many related series with rich covariates. On short, clean, stable series, a well-tuned statistical model is competitive and far cheaper to run.
It depends on SKU count, not company size. A small business with a handful of stable products is genuinely better off with a spreadsheet and a naive baseline. The case for AI shows up once you’re managing hundreds of SKUs, frequent promotions, or several new-product launches a year.
Open-source, it’s Python libraries: statsmodels for ARIMA and ETS, LightGBM or XGBoost for gradient boosting, and libraries like Chronos or TimesFM for foundation-model baselines. Commercial supply-chain platforms also exist, but verify any accuracy claim against your own backtest, not the vendor’s.
No. It replaces the manual averaging and ad hoc adjustment. The planner’s job shifts toward setting the decision the forecast serves, reading the model’s reasoning, and overriding it only when they have a genuinely good reason to, backed by Forecast Value Added, not instinct.
Not from its own history, since there isn’t any yet. Teams forecast by analogy to similar past launches, or use a pre-trained time-series foundation model as a starting point, then correct quickly once early sales come in.
There’s no universal number. It depends on volume, volatility and horizon. The meaningful test is whether the forecast beats a naive baseline, that’s Forecast Value Added.
It depends on data readiness, not modelling. A first backtested baseline can take a matter of weeks. Getting clean, reconciled, production-ready data underneath it usually takes considerably longer.





