Your sales history is not demand history. It is demand filtered through every promo you ran, every stockout you had, and every price you set, and a model trained on it raw learns your past decisions instead of your customers. The fix is not a smarter model. It is cleaning the lies out of the series before any model sees it.
October planning meeting, every retailer, every year. A planner pulls last November's sales by SKU to seed the holiday forecast, and last November contains a 30%-off sitewide flash, a hero SKU that stocked out on day 2 of Black Friday week, and a category that did not exist until March.
The model will fit that series beautifully. It will learn that the hero SKU sells 4 units a day in peak week, because that is what the data says happened, and the data is not lying about what was recorded. It is lying about what was demanded.
Holiday forecasting fails less often on algorithm choice than on this gap between recorded sales and actual demand. So before arguing about gradient boosting versus exponential smoothing, walk the 3 ways the history lies and the correction for each.
Lie 1: promotions are baked into the baseline
A promo week shows demand at a price that will not repeat. Train on it unlabeled and the model splits the uplift between seasonality and noise, which corrupts both: your November seasonal index absorbs the discount you happened to run, and next year's forecast silently assumes you will run it again.
The correction is to model promos as explicit inputs rather than scrub them out. Prophet's holiday and regressor interface is the clearest teaching example of the shape: you hand the model a dated list of events, each with a window of days before and after, and it fits each event's effect separately from trend and seasonality.
The window parameters matter more than they look. A flash sale pulls demand forward, so the days after it run below baseline, and an event modeled without its trailing window pushes that dip into the seasonal component instead.
Whatever library you use, the discipline is the same and it is organizational before it is statistical. The promo calendar has to exist as structured data, past and planned, with dates, depth, and scope. A forecasting team that has to reconstruct last year's promos from finance emails has found its real project.
Lie 2: stockouts record zero sales, not zero demand
The day the hero SKU sold out, recorded sales fell to 0 while demand kept arriving. Statisticians call this censoring: the true value existed and the measurement was capped at what inventory allowed.
Feed censored zeros to a model and it learns that demand collapses right after your best days. The result is the cruelest failure in retail forecasting, where last year's stockout becomes this year's under-buy, which causes this year's stockout, which trains next year's model. Teams live inside that loop for years and read it as customers being unpredictable.
The correction has 2 parts, one cheap and one honest.
- Cheap: join sales history to inventory history and flag every period where the item was unavailable or nearly so. Treat flagged periods as missing data rather than as zeros, and let the model interpolate from surrounding demand.
- Honest: reconstruct the censored demand where you can, from waitlist signups, search volume for the product, or the sales velocity of the hours before the stockout, and record the estimate with a flag saying it is one.
Either beats the default. The unforgivable version is the one most pipelines run: zeros in, no flags, and a model quietly learning that scarcity is a demand signal.
Every observation in the training window needs an answer to one question: could the customer actually have bought this, at what price, in what context? A row that cannot answer is not evidence. It is noise wearing evidence's clothes.
Lie 3: the new SKU has no history at all
A third of a typical holiday assortment is new: fresh colorways, gift sets that exist for 8 weeks, this year's bundle pricing. Per-SKU time-series models have nothing to fit, and the common workaround of copying last year's nearest item is a guess with extra steps.
Two mechanisms do better, and they compose.
- Borrow through attributes. Forecast at the level where history exists, category by price band by channel, then allocate down to new SKUs by attribute similarity to items that have sold. The new fleece in a proven category inherits the category's curve, scaled by its price position.
- Respect intermittency. Gift-shop items that sell in bursts with long gaps between break standard smoothing methods. Croston's method exists for exactly this pattern, splitting the series into demand size and interval between demands, and the statsforecast implementation makes it a 3-line baseline rather than a research project.
Cold-start forecasts deserve their own error tracking, separate from established SKUs. Averaging them together lets a strong catalog hide a weak new-item process, and new items are where the holiday over-buy and under-buy money actually sits.
The hierarchy problem: your forecasts disagree with each other
Forecast every SKU independently and sum the results, and the sum will not match a forecast made directly at category level. Both numbers land in different meetings, the buying team plans against one and finance against the other, and the argument that follows has no resolution because both are outputs of the same pipeline.
This is the reconciliation problem, and it has real methods rather than a meeting. The hierarchical forecasting chapters of Forecasting: Principles and Practice lay out the options: bottom-up summing, top-down allocation, and reconciliation approaches that adjust forecasts at every level to be coherent, so that lower levels genuinely sum to upper ones.
The practical takeaway is not which method wins. It is that coherence is a constraint you impose deliberately, with a documented choice, instead of a property you assume and discover missing during peak-week allocation. The hierarchicalforecast library implements the standard reconciliation methods against forecasts you already produce, which makes the experiment cheap.
Reconciliation is a governance tool as much as a statistical one. One coherent set of numbers means buying, finance and stores plan against the same future, and the argument moves from dueling spreadsheets to a documented method choice.
Retail hierarchies are also plural, which is the trap inside the trap. Product rolls up SKU to style to category, geography rolls up store to region, and both hierarchies must hold at once. Pick the crossed structure explicitly before writing pipeline code, because retrofitting a second hierarchy into a reconciliation step is a rewrite, not a patch.
What the M5 competition settles, and what it cannot
For calibration on methods, the M5 forecasting competition is the best public reference point retail has. It ran on hierarchical Walmart unit-sales data at item, department, category and store level across 3 US states, with explanatory variables including price, promotions and special events, and it scored accuracy across the levels of the hierarchy rather than at one.
The competition's setup is itself the lesson. The organizers included prices and events as inputs because item-level retail demand is not forecastable from its own past alone, and they scored hierarchically because a forecast that is right in aggregate and wrong at store-item level does not ship boxes to the right place.
What no competition settles is your data quality. A method that ranks well on a curated benchmark still inherits every uncorrected stockout and unlabeled promo in your series, which is why the corrections above come before model selection rather than after.
Backtest the way you eval
Holiday models get validated on ordinary weeks because that is what most of the data is, and then they meet the 6 weeks that break every assumption. The backtest has to be seasonal to mean anything: hold out last year's October-through-December, train on everything before it, and score only the held-out peak.
Run it per-slice, exactly like a golden set with a per-slice floor. Accuracy on established SKUs versus new, promoted versus not, stable stores versus new ones. An aggregate error number hides the slice that costs money, and in holiday forecasting the expensive slice is nearly always promoted new items.
Score against corrected demand, not recorded sales, for any held-out period with a stockout. Grading a forecast as wrong for predicting demand the shelf could not supply teaches the pipeline the same lie you just cleaned out of training.
When the human should override the model
Some of what makes this November unlike last November exists only in people's heads: a competitor's collapse, a port delay pushing your top category to late arrival, a marketing budget that doubled. No amount of history correction encodes information that is not in the history yet.
Judgmental adjustment has an evidence base, and it is more skeptical than either camp expects. The judgmental forecasting chapter of Forecasting: Principles and Practice documents both sides: judgment is subject to anchoring, optimism and political pressure, and it also improves forecasts when it brings information the model lacks, applied through a structured process rather than a hallway conversation.
The structure that keeps overrides honest fits in 4 rules.
- Overrides go on the model's number, recorded as a delta with a stated reason, never as a replacement that erases what the model said.
- The reason must name information the model does not have. Disliking the number is not information.
- Big adjustments get a second signature, the same way big refunds do.
- After the season, score the overrides against the model's originals, per person, and publish the result internally.
That last rule changes behavior more than the other 3 combined. Planners whose adjustments beat the model earn wider bounds next year, and the ones whose adjustments consistently cost accuracy start writing better reasons or fewer overrides.
The corrections, in one table
| How history lies | What the model learns | Correction |
|---|---|---|
| Promo weeks unlabeled | Discount uplift becomes seasonality | Promo calendar as regressors with pre and post windows |
| Stockout zeros | Demand collapses after best days | Censor-flag from inventory joins; treat as missing, not zero |
| Pull-forward after events | Post-event dip reads as weak weeks | Model trailing windows on every event, beyond the day itself |
| New SKUs with no series | Nothing; the model guesses or copies | Attribute-based borrowing from the level where history exists |
| Intermittent gift items | Smoothed averages that miss the bursts | Intermittent-demand methods like Croston as the baseline |
| Incoherent hierarchy | SKU forecasts that contradict category plans | Explicit reconciliation across product and location at once |
| Ordinary-week validation | Confidence that evaporates in November | Seasonal holdout scored per-slice against corrected demand |
None of these corrections requires a new platform, and most are joins and flags against data already sitting in the order and inventory systems. The work is unglamorous, which is why it loses planning-cycle arguments to model upgrades that matter less.
The correction pipeline: label the lies before any model fits them, and route human knowledge through scored overrides, not edits.
FAQ
Should we exclude last year's Black Friday from training as an outlier?
No. Deleting it removes your only observation of peak behavior. Label it, with its promo depth and stockout flags, so the model learns peak conditioned on what you did rather than peak as random noise.
How much history do we need for a usable holiday forecast?
Two full holiday cycles is the practical floor for seasonal structure at item level, and 3 is comfortable. With less, forecast higher in the hierarchy where more cycles exist and allocate down, which is the cold-start mechanism applied to the whole calendar.
Is a machine learning model worth it over exponential smoothing for us?
Run both against the seasonal holdout and let the backtest answer. The corrections in this post move error more than the model family for most mid-size retailers, and a corrected series makes every model on top of it better.
Who should own demand-history correction?
The forecasting team owns the code, but the promo calendar needs an owner in marketing and the stockout flags need one in inventory operations. Corrections rot when the data they join against has no owner.
What accuracy should we promise the buying team?
Promise a measured error range per level of the hierarchy from the backtest, not a single number. Item-week forecasts carry error that looks alarming out of context, and aggregates are far tighter, so pair each planning decision with the level whose error supports it.
Do judgmental overrides mean the model failed?
An override with new information is the system working as designed. A pattern of overrides repeating the same reason means that reason should become a model input, which is how the override log turns into next year's feature list.
References
- Prophet - seasonality, holiday effects, and regressors
- Hyndman & Athanasopoulos, Forecasting: Principles and Practice - forecasting hierarchical and grouped time series
- Hyndman & Athanasopoulos, Forecasting: Principles and Practice - judgmental forecasts
- Nixtla - hierarchicalforecast reconciliation library
- Nixtla statsforecast - Croston's method for intermittent demand
- Kaggle - M5 forecasting accuracy competition
Related reading
More in AIThe Human in the Loop Is a Role, Not a Checkbox
A working loop is a staffed role: a named reviewer with domain ownership, a review surface built for verification, an explicit split between gating and sampling, triggers that tighten and loosen it, and a verdict wire back into the eval suite.
Evals Before Agents: The Regression Suite Is What Makes an AI Feature Shippable
An AI feature without a regression suite is a demo with a deploy pipeline. What a golden set looks like for real retail workflows, the 4 grader types in cost order, the 3 kinds of drift, and where the human approval gate belongs.
Why AI Pilots Die Before Production
The demo works, the stakeholders are pleased, and 7 months later it is quietly switched off. The model is almost never the reason. Five failure modes account for most of it: no evals, no data contracts, no cost ceiling, guardrails scoped too late, and no approval path.