Skip to main content
Back to AI Commerce Lab
AI·October 2026·10 min read

Forecasting Holiday Demand When Your History Lies

Your sales history is not demand history. It is demand filtered through every promo you ran, every stockout you had, and every price you set, and a model trained on it raw learns your past decisions instead of your customers. The fix is not a smarter model. It is cleaning the lies out of the series before any model sees it.

October planning meeting, every retailer, every year. A planner pulls last November's sales by SKU to seed the holiday forecast, and last November contains a 30%-off sitewide flash, a hero SKU that stocked out on day 2 of Black Friday week, and a category that did not exist until March.

The model will fit that series beautifully. It will learn that the hero SKU sells 4 units a day in peak week, because that is what the data says happened, and the data is not lying about what was recorded. It is lying about what was demanded.

Holiday forecasting fails less often on algorithm choice than on this gap between recorded sales and actual demand. So before arguing about gradient boosting versus exponential smoothing, walk the 3 ways the history lies and the correction for each.

Lie 1: promotions are baked into the baseline

A promo week shows demand at a price that will not repeat. Train on it unlabeled and the model splits the uplift between seasonality and noise, which corrupts both: your November seasonal index absorbs the discount you happened to run, and next year's forecast silently assumes you will run it again.

The correction is to model promos as explicit inputs rather than scrub them out. Prophet's holiday and regressor interface is the clearest teaching example of the shape: you hand the model a dated list of events, each with a window of days before and after, and it fits each event's effect separately from trend and seasonality.

The window parameters matter more than they look. A flash sale pulls demand forward, so the days after it run below baseline, and an event modeled without its trailing window pushes that dip into the seasonal component instead.

Whatever library you use, the discipline is the same and it is organizational before it is statistical. The promo calendar has to exist as structured data, past and planned, with dates, depth, and scope. A forecasting team that has to reconstruct last year's promos from finance emails has found its real project.

Lie 2: stockouts record zero sales, not zero demand

The day the hero SKU sold out, recorded sales fell to 0 while demand kept arriving. Statisticians call this censoring: the true value existed and the measurement was capped at what inventory allowed.

Feed censored zeros to a model and it learns that demand collapses right after your best days. The result is the cruelest failure in retail forecasting, where last year's stockout becomes this year's under-buy, which causes this year's stockout, which trains next year's model. Teams live inside that loop for years and read it as customers being unpredictable.

The correction has 2 parts, one cheap and one honest.

  • Cheap: join sales history to inventory history and flag every period where the item was unavailable or nearly so. Treat flagged periods as missing data rather than as zeros, and let the model interpolate from surrounding demand.
  • Honest: reconstruct the censored demand where you can, from waitlist signups, search volume for the product, or the sales velocity of the hours before the stockout, and record the estimate with a flag saying it is one.

Either beats the default. The unforgivable version is the one most pipelines run: zeros in, no flags, and a model quietly learning that scarcity is a demand signal.

Every observation in the training window needs an answer to one question: could the customer actually have bought this, at what price, in what context? A row that cannot answer is not evidence. It is noise wearing evidence's clothes.

Lie 3: the new SKU has no history at all

A third of a typical holiday assortment is new: fresh colorways, gift sets that exist for 8 weeks, this year's bundle pricing. Per-SKU time-series models have nothing to fit, and the common workaround of copying last year's nearest item is a guess with extra steps.

Two mechanisms do better, and they compose.

  1. Borrow through attributes. Forecast at the level where history exists, category by price band by channel, then allocate down to new SKUs by attribute similarity to items that have sold. The new fleece in a proven category inherits the category's curve, scaled by its price position.
  2. Respect intermittency. Gift-shop items that sell in bursts with long gaps between break standard smoothing methods. Croston's method exists for exactly this pattern, splitting the series into demand size and interval between demands, and the statsforecast implementation makes it a 3-line baseline rather than a research project.

Cold-start forecasts deserve their own error tracking, separate from established SKUs. Averaging them together lets a strong catalog hide a weak new-item process, and new items are where the holiday over-buy and under-buy money actually sits.

The hierarchy problem: your forecasts disagree with each other

Forecast every SKU independently and sum the results, and the sum will not match a forecast made directly at category level. Both numbers land in different meetings, the buying team plans against one and finance against the other, and the argument that follows has no resolution because both are outputs of the same pipeline.

This is the reconciliation problem, and it has real methods rather than a meeting. The hierarchical forecasting chapters of Forecasting: Principles and Practice lay out the options: bottom-up summing, top-down allocation, and reconciliation approaches that adjust forecasts at every level to be coherent, so that lower levels genuinely sum to upper ones.

The practical takeaway is not which method wins. It is that coherence is a constraint you impose deliberately, with a documented choice, instead of a property you assume and discover missing during peak-week allocation. The hierarchicalforecast library implements the standard reconciliation methods against forecasts you already produce, which makes the experiment cheap.

Reconciliation is a governance tool as much as a statistical one. One coherent set of numbers means buying, finance and stores plan against the same future, and the argument moves from dueling spreadsheets to a documented method choice.

Retail hierarchies are also plural, which is the trap inside the trap. Product rolls up SKU to style to category, geography rolls up store to region, and both hierarchies must hold at once. Pick the crossed structure explicitly before writing pipeline code, because retrofitting a second hierarchy into a reconciliation step is a rewrite, not a patch.

What the M5 competition settles, and what it cannot

For calibration on methods, the M5 forecasting competition is the best public reference point retail has. It ran on hierarchical Walmart unit-sales data at item, department, category and store level across 3 US states, with explanatory variables including price, promotions and special events, and it scored accuracy across the levels of the hierarchy rather than at one.

The competition's setup is itself the lesson. The organizers included prices and events as inputs because item-level retail demand is not forecastable from its own past alone, and they scored hierarchically because a forecast that is right in aggregate and wrong at store-item level does not ship boxes to the right place.

What no competition settles is your data quality. A method that ranks well on a curated benchmark still inherits every uncorrected stockout and unlabeled promo in your series, which is why the corrections above come before model selection rather than after.

Backtest the way you eval

Holiday models get validated on ordinary weeks because that is what most of the data is, and then they meet the 6 weeks that break every assumption. The backtest has to be seasonal to mean anything: hold out last year's October-through-December, train on everything before it, and score only the held-out peak.

Run it per-slice, exactly like a golden set with a per-slice floor. Accuracy on established SKUs versus new, promoted versus not, stable stores versus new ones. An aggregate error number hides the slice that costs money, and in holiday forecasting the expensive slice is nearly always promoted new items.

Score against corrected demand, not recorded sales, for any held-out period with a stockout. Grading a forecast as wrong for predicting demand the shelf could not supply teaches the pipeline the same lie you just cleaned out of training.

When the human should override the model

Some of what makes this November unlike last November exists only in people's heads: a competitor's collapse, a port delay pushing your top category to late arrival, a marketing budget that doubled. No amount of history correction encodes information that is not in the history yet.

Judgmental adjustment has an evidence base, and it is more skeptical than either camp expects. The judgmental forecasting chapter of Forecasting: Principles and Practice documents both sides: judgment is subject to anchoring, optimism and political pressure, and it also improves forecasts when it brings information the model lacks, applied through a structured process rather than a hallway conversation.

The structure that keeps overrides honest fits in 4 rules.

  • Overrides go on the model's number, recorded as a delta with a stated reason, never as a replacement that erases what the model said.
  • The reason must name information the model does not have. Disliking the number is not information.
  • Big adjustments get a second signature, the same way big refunds do.
  • After the season, score the overrides against the model's originals, per person, and publish the result internally.

That last rule changes behavior more than the other 3 combined. Planners whose adjustments beat the model earn wider bounds next year, and the ones whose adjustments consistently cost accuracy start writing better reasons or fewer overrides.

The corrections, in one table

How history liesWhat the model learnsCorrection
Promo weeks unlabeledDiscount uplift becomes seasonalityPromo calendar as regressors with pre and post windows
Stockout zerosDemand collapses after best daysCensor-flag from inventory joins; treat as missing, not zero
Pull-forward after eventsPost-event dip reads as weak weeksModel trailing windows on every event, beyond the day itself
New SKUs with no seriesNothing; the model guesses or copiesAttribute-based borrowing from the level where history exists
Intermittent gift itemsSmoothed averages that miss the burstsIntermittent-demand methods like Croston as the baseline
Incoherent hierarchySKU forecasts that contradict category plansExplicit reconciliation across product and location at once
Ordinary-week validationConfidence that evaporates in NovemberSeasonal holdout scored per-slice against corrected demand

None of these corrections requires a new platform, and most are joins and flags against data already sitting in the order and inventory systems. The work is unglamorous, which is why it loses planning-cycle arguments to model upgrades that matter less.

RAW HISTORY promos baked in stockout zeros new-SKU gaps CORRECTIONS promo regressors censor flags reconciliation FORECAST scored per-slice JUDGMENTAL OVERRIDE delta + reason scored overrides feed next year's inputs

The correction pipeline: label the lies before any model fits them, and route human knowledge through scored overrides, not edits.

FAQ

Should we exclude last year's Black Friday from training as an outlier?

No. Deleting it removes your only observation of peak behavior. Label it, with its promo depth and stockout flags, so the model learns peak conditioned on what you did rather than peak as random noise.

How much history do we need for a usable holiday forecast?

Two full holiday cycles is the practical floor for seasonal structure at item level, and 3 is comfortable. With less, forecast higher in the hierarchy where more cycles exist and allocate down, which is the cold-start mechanism applied to the whole calendar.

Is a machine learning model worth it over exponential smoothing for us?

Run both against the seasonal holdout and let the backtest answer. The corrections in this post move error more than the model family for most mid-size retailers, and a corrected series makes every model on top of it better.

Who should own demand-history correction?

The forecasting team owns the code, but the promo calendar needs an owner in marketing and the stockout flags need one in inventory operations. Corrections rot when the data they join against has no owner.

What accuracy should we promise the buying team?

Promise a measured error range per level of the hierarchy from the backtest, not a single number. Item-week forecasts carry error that looks alarming out of context, and aggregates are far tighter, so pair each planning decision with the level whose error supports it.

Do judgmental overrides mean the model failed?

An override with new information is the system working as designed. A pattern of overrides repeating the same reason means that reason should become a model input, which is how the override log turns into next year's feature list.

References

From the Destm engineering archive. For current work on this topic, start at Solutions or the blog.