← All articles
Case Study7 min read

Demand Forecasting in Retail: Build the Boring Baseline Before the Model

An illustrative retail scenario showing how we approach AI demand forecasting: simple baselines, catalogue segmentation, backtests that mimic real ordering, and a human review loop buyers can trust.

SN
Stoyan Nikolov
Software Architect · 7 Oct 2026

Demand forecasting is the AI project that sounds easiest to justify and is easiest to get quietly wrong. Every retailer has sales history, every manager has an opinion about overstock, and the pitch writes itself: predict what will sell, order accordingly, waste less. We have been involved in enough of these conversations to say that the model is rarely where the project succeeds or fails. The decisions around it are.

Everything below is an illustrative scenario, not a client story. We use it because the pattern repeats across the shops and distributors we talk to.

The scenario: a regional shop with 4,000 SKUs

Imagine a regional retailer with around forty stores and a mixed catalogue: fast-moving groceries, seasonal garden goods, and a long tail of items that sell twice a month. Replenishment today is a mix of spreadsheet rules and the instincts of two experienced buyers. The owner wants "AI forecasting".

The first instinct is to feed all 4,000 products into one model and compare its output with the buyers. We would resist that. A single accuracy number across such a catalogue hides the thing that matters, because the fast movers dominate any aggregate score while the long tail is where the guessing hurts.

Start with a baseline that is embarrassingly simple

Before any machine learning, we build the dumbest forecast that could work: last year's same-week sales, or a moving average with a seasonal adjustment. This is not a formality. The standard open textbook on the subject, Forecasting: Principles and Practice by Hyndman and Athanasopoulos, treats simple benchmark methods as the yardstick every more complex method must beat. If your model cannot clearly beat the seasonal-naive baseline on the products you care about, you do not have a model problem to solve; you have a data or process problem.

It is also a useful honesty check on the buyers. Sometimes the experienced humans are beating the baseline by a wide margin because they know about a supplier delay or a local event that never appears in the data. That knowledge is an input to capture, not a competitor to replace.

Segment before you model

We split the catalogue into three groups and treat them differently.

This is the first place the "which processes should not use AI" question becomes concrete. Excluding the third group from the model is not a failure of ambition. It keeps the project's accuracy figures honest and keeps effort where money actually moves.

What the research does and does not tell you

The largest public benchmark in retail forecasting is the M5 competition, built on Walmart sales data. Its published results, summarised in the M5 accuracy competition paper, showed that machine-learning approaches generally outperformed the statistical benchmarks on that dataset. That is a real and useful finding. It is not a promise about your data. Walmart's volume, history, and hierarchy are not a regional chain's, and a competition setting rewards a single error metric in a way a business does not. We read it as evidence that modern methods are worth testing, not as a reason to skip the baseline.

Measure in money, not only in percentages

Forecast error metrics are necessary and insufficient. A two-percent improvement means little until you translate it into stock decisions. Over-forecasting a cheap, durable item costs shelf space. Over-forecasting fresh produce costs the product itself. Under-forecasting a high-margin item costs the sale and possibly the customer.

So we ask the business to state, per segment, which mistake is worse. That asymmetry shapes the model's objective and the safety stock. In practice it is a short, slightly uncomfortable workshop, and it produces more value than any hyperparameter search we have run.

The evaluation we recommend is a backtest that mimics real ordering. Pick a historical date, forecast only with data available then, place the orders the system would have placed, and compare resulting stock-outs and write-offs against what actually happened. Random train/test splits flatter models because they leak information from the future, and a forecast that looks excellent in a notebook can be useless on a Monday morning.

Design the human loop on day one

The forecast should arrive as a proposed order, not as a number in a dashboard. Buyers see the suggestion, the baseline, and the main reasons: a promotion starting, a seasonal peak, an unusual recent week. They can override, and the override is recorded with a short reason. Two things follow from that. Trust grows because people can see why, and the overrides become the best training signal and bug report we have. If buyers keep correcting the same category, the system is missing an input.

We also set explicit rules for when the system should not act on its own: a product with a supplier lead-time change, a sudden sales spike with no known cause, a data feed that arrived late or incomplete. Silent failure in an upstream feed is the most common way we have seen forecasts go bad, and a model will happily produce confident numbers from broken data.

Where it tends to go wrong

Stock-outs censor demand. If an item sold out on Tuesday, recorded sales understate what customers wanted, and a model trained on that learns to under-order again. Handling this means flagging availability in the data and treating those days differently, which is dull work that decides whether results are credible.

Another trap is changing the process and the model at the same time. If the retailer also shortens supplier cycles during the pilot, nobody can say which change produced the improvement. We prefer a staged rollout: a handful of stores or categories on the new approach, the rest as control, run long enough to cover at least one meaningful seasonal swing.

What we would do first

If this retailer came to us, the first four weeks would involve very little modelling. We would clean the sales and availability history, agree the promotion calendar, build the baseline, run the backtest, and write down with the buyers which errors cost what. Only then would we decide whether a more complex model is justified, and for which segment.

That order feels slow to people who want to see AI working. It is faster than the alternative, which is a polished forecast nobody trusts and nobody can evaluate. A forecasting project is ready for production when the team can explain, in plain language and in euros, why the system's orders beat the old ones. If you cannot say that yet, keep working on the baseline.

Have a project in mind?

Tell us what you want to build and we’ll come back with a plan.

Start a project