Blog

Three a Day Is Harder to Forecast Than Three Hundred

Sep 29
decor image

By Matt Wampler, CEO of ClearCOGS

An operator running a growing group in a smaller market described her problem to me in one sentence that I thought was unusually precise. She said her stores get plenty of tickets, just not as many as the brand’s larger markets, and because of that nobody can work out how much product to have ready at any given hour.

She had identified something real, and she had identified it more clearly than most people do. The issue was not that her team was careless or that her data was bad. It was that she was asking for a number at a level of detail her volume could not support.

This comes up constantly, and it is worth understanding, because the usual response makes it worse.

Small counts are not small versions of big counts

Start with the arithmetic, because the intuition is unhelpful here.

A location selling three hundred units of something across a day has a demand pattern that is fairly stable. Day-to-day swings exist, but they are modest relative to the total. Predict three hundred, land within thirty, and you are close enough to operate on.

Now take the same item at a store or an hour where the expected count is three. The absolute swing is tiny, but the relative swing is enormous. Real days will land on one, or five, or zero. Nothing is wrong with the forecast. That is simply how small counts behave, and no amount of history changes it, because the variation is in the demand itself rather than in your estimate of it.

Here is the practical consequence, with illustrative figures:

Expected units in the intervalA typical realistic rangeWhat that means operationally
300Roughly 265 to 335A forecast is directly actionable
50Roughly 36 to 64Useful, with a buffer on the short side
10Roughly 4 to 16Directionally useful, not a production target
3Roughly 0 to 7A point forecast is close to meaningless

Illustrative only. These ranges show how natural variation scales with volume; they are not benchmarks.

Every row is the same underlying stability. What changes is how much of the total the natural variation represents. A production number is useful when the expected value dominates the noise, and stops being useful when it does not.

The trap operators fall into

When the hourly number looks unreliable, the instinct is to ask for more precision. Break the day into smaller intervals. Forecast the ingredient rather than the item. Get it per station.

Every one of those moves makes the underlying problem worse, because each one reduces the number of observations feeding the estimate. You end up with a more detailed number that is less trustworthy, delivered to a manager who now has evidence that the system does not work.

That is the real cost. Not the inaccuracy itself, which was unavoidable at that grain, but the loss of credibility that follows. A team that gets a wrong hourly number for a slow period will stop reading the daily number too, and the daily number was fine.

Not every item is the same forecasting problem

The forecasting literature figured this out decades ago in a different industry, and the framing transfers cleanly.

Researchers studying demand patterns argue that you have to classify a series before you choose a method for it, and they note that the common practice in inventory control software is to categorize those demand patterns arbitrarily and then proceed to pick an estimation procedure anyway. Their categorization sorts series along two dimensions: how often demand occurs at all, and how much the size of demand varies when it does. That yields four groups, described as smooth, erratic, intermittent, and lumpy, with different methods suited to each. The rules were tested against 3,000 real demand series (Syntetos, Boylan, and Croston, Journal of the Operational Research Society, 2005).

The underlying claim is the useful part for an operator: a single forecasting approach applied uniformly across a menu will do well on the items that happen to suit it and badly on the rest. Your best seller and your slowest side dish are not the same statistical problem, and treating them as one is why the long tail of the menu always looks like the forecast is broken.

Where this shows up in a restaurant

Four places, and the profile is recognizable in each.

Slow dayparts. The late afternoon between lunch and dinner. The last hour before close. These intervals often hold a handful of transactions, and they are exactly where hourly guidance gets requested most, because that is where waste feels worst.

Small-format and new locations. A store in a modest trade area, or one open eight weeks, has thin history and low counts at once. Holding it to the same forecast granularity as a flagship guarantees it looks worse.

The long tail of the menu. In most brands, a small number of items drive the majority of volume, and a much larger number sell in ones and twos. The tail is where item-level forecasting appears to fail, and it is also where the money involved is smallest.

Ingredient-level detail on low-volume items. Forecasting an ingredient used only in a slow-selling dish inherits that dish’s noise and adds recipe assumptions on top.

What to do instead

Four adjustments, in rough order of impact.

  1. Match the grain to the volume. Set forecast granularity by how much the interval actually moves, not by a uniform brand standard. Busy items at a busy store can support hourly targets. A slow afternoon may only support a daypart number, and a long-tail item may only support a weekly one. Different rows in the same report can legitimately run at different grains.
  2. Pool across locations for the tail. A slow item at one store is noisy. The same item across thirty stores is a usable pattern. Estimate the shape from the group and scale it to each location rather than trying to learn it thirty times independently from almost no data.
  3. Give low-volume items a rule, not a forecast. For items where the expected count is small, a minimum hold plus make-to-order is usually a better operating answer than a predicted quantity. That is not a failure to forecast. It is recognizing which decisions a forecast can improve and which it cannot.
  4. Say which numbers are firm and which are directional. Managers handle uncertainty well when it is labeled. The damage comes from presenting a high-confidence daily number and a low-confidence hourly one in identical formatting, which teaches people to distrust both equally.

More detail is not the same thing as more accuracy. Past a certain point they move in opposite directions, and the crossover happens at the volume of the interval you are forecasting, not at the quality of your model.

The operator I mentioned had the diagnosis right. What she needed was not a better hourly number. It was to find the level at which her volume makes a number meaningful, operate on that, and stop paying an accuracy penalty for detail nobody could act on anyway.

This is the work we spend our days on at ClearCOGS: deciding what grain each item at each location can actually support, and delivering a number a manager can trust at that grain rather than a precise-looking one they cannot.

If parts of your menu or parts of your day look unforecastable, that is usually a question about volume rather than about the model.

Let’s Talk

Sources

  • Syntetos, A. A., Boylan, J. E., and Croston, J. D. On the Categorization of Demand Patterns. Journal of the Operational Research Society, 56(5), 495–503. May 2005. link.springer.com