By Matt Wampler, CEO of ClearCOGS
Quick Answer
A four-week average takes the same weekday from the last four weeks and averages it to estimate today. It is the default prep method in most point of sale systems because it is simple and requires no setup. It fails because it assumes recent same-weekdays are interchangeable, carries no information about why demand moved, and becomes least reliable exactly where prep decisions are tightest: at small counts, in short time windows, and at lower volume locations.
What Is a Four-Week Average in Restaurant Prep?
A four-week average, sometimes called a trailing or rolling average, estimates demand for a given day or time slot by averaging the same slot from the previous four weeks.
If a store wants to know how much bone-in chicken to hold at 6:00 p.m. on a Friday, the system looks at the last four Fridays at 6:00 p.m., adds those numbers, and divides by four. Many systems will do this at a 30 minute granularity across a whole day and print the result on a prep sheet.
It is worth being fair to the method. A trailing average is a legitimate statistical benchmark. It is cheap, it is transparent, and a manager can verify it by hand. For a stable, high volume, low variance operation, it is often close enough that nobody notices what it is missing.
The trouble starts when an operation is not stable, not high volume, or not low variance. Which describes most restaurants.
Why Does the Four-Week Average Fail?
It Treats Every Same-Weekday as Interchangeable
The method’s core assumption is that last Friday, the Friday before that, and the two before those are all reasonable stand-ins for this Friday. They usually are not. One of them had a promotion running. One had rain through the dinner window. One followed a local event. One was the first Friday of the month, which in many markets behaves differently from the third.
The average does not know any of this. It smooths four different Fridays into one number and hands you the middle, which is a value that describes none of the four days and may describe today even less.
It Cannot Explain Itself, So It Cannot Adapt
An average has no mechanism for knowing why demand moved. That has a practical consequence: it cannot respond to anything it has not already absorbed.
When a limited time offer launches, the average takes four weeks to catch up, and then takes another four weeks to let go of it after the promotion ends. Every prep decision in those windows is systematically wrong in a known direction, and the method has no way to tell you which direction.
The same is true of weather, school calendars, road construction, and a competitor opening down the street. These are not exotic variables. They are the ordinary reasons a Friday stops looking like a Friday.
Its Reliability Depends on How Large the Counts Are
This is the failure operators feel most and name least, because it is a property of the arithmetic rather than the restaurant.
The standard reference on forecasting method, Hyndman and Athanasopoulos, makes the point directly in its section on time series of counts. Most common forecasting methods assume the data lives on a continuous scale, but a great deal of real demand data arrives as counts, and you cannot serve 3.45 customers. The authors note that this distinction stops mattering once counts are sufficiently large, using a minimum of roughly 100 as the point where the difference has no perceivable effect. Below that, with small counts, they are explicit that different methods are required.
Read that against a prep sheet. A store’s total daily transactions might be a large, well-behaved number. But a prep sheet does not ask about daily totals. It asks how many pieces of one item to have ready in one 30 minute window. Slice a day into 30 minute buckets and split it across a dozen menu items, and the number in each cell is small.
The averaging method that behaves acceptably on the daily total is being applied to cells where it was never appropriate. The result is not a small error. It is an unstable one, moving in different directions on different days for no reason the manager can see, which is why teams stop trusting the sheet and go back to eyeballing it.
Why Do Lower Volume Locations Get the Worst Results?
This is the counterintuitive part, and it matters because it inverts a common assumption.
Operators in smaller markets often assume forecasting is a large-chain concern. The reasoning goes that big systems have the volume to justify the investment, while a smaller group should stick to simple tools until it grows.
The arithmetic points the other way. Because averaging degrades as counts get smaller, a lower volume location is exactly where a trailing average performs worst. A store doing high volume has enough transactions in every window that the noise partly cancels out. A store doing modest volume in the same window does not. Its four-week average swings, and the manager compensating for those swings usually compensates upward, because running out is more visible than throwing away.
So the smaller operation gets the least reliable number and pays for it with the thinner margin. It is precisely backwards from how the decision usually gets made.
Growth makes this sharper rather than softer. An operator adding locations quickly has stores with only a few weeks of history, where a four-week window is either incomplete or dominated by opening-period behavior that will not repeat.
Why Is the Hourly Prep Number Harder Than the Daily One?
Because for held product, the daily number is not a decision. It is a summary.
If a product has a hold window measured in minutes, the operational question is not how much to make today. It is how much to have ready at 11:30, and again at 12:00, and again at 12:30. Get the total right and the distribution wrong, and you have simultaneously run out at noon and thrown away product at two o’clock. The daily report will show a perfectly reasonable number and reveal nothing about either failure.
This is why intraday granularity is not a refinement of prep planning. For a fresh or held product, it is the actual unit of the problem. And it is the granularity where small counts bite hardest, which is why the two failures reinforce each other.
Rolling Average Versus Demand Forecast
| Dimension | Trailing four-week average | Demand forecast |
|---|---|---|
| Inputs | Same weekday, last four weeks | Full item-level sales history, plus calendar, promotions, weather, local events |
| Handles promotions | No, lags in and lags out | Yes, treated as a known input |
| Behavior at small counts | Unstable, error grows as counts shrink | Method selected to suit the count size |
| Time granularity | Whatever the system defaults to | Chosen to match the product’s hold window |
| New locations | Needs four weeks of clean history | Can borrow patterns from comparable locations |
| Accuracy measurable | Rarely tracked | Predicted versus actual, item by item, day by day |
The last row is the one that changes how a team behaves. An average is almost never scored, so nobody knows whether it is working. A forecast that publishes what it predicted next to what actually happened can be held accountable, and a number that can be checked is a number a kitchen manager will eventually use.
What Does a Prep Forecast Need Instead?
Less than most operators expect, and it does not include a finished back office cleanup.
- Item-level transaction history from the POS. What sold, when, at what granularity. This is usually the cleanest data a restaurant owns, because nobody types it in by hand.
- Recipes with yields and pack sizes. These do not need to be in a particular system. A recipe book or a spreadsheet works.
- The real hold window for each item, which is often shorter than the documented shelf life and is frequently the constraint that matters most.
- Calendar and promotional context, so the model treats a limited time offer as information rather than noise.
- Local conditions, meaning weather and events near each specific location, not regional averages.
- An accuracy record, so the team can see predicted against actual and decide for themselves whether to trust it.
Frequently Asked Questions
Is a rolling average a forecast?
Not really. It is a benchmark. In forecasting practice, simple methods like a trailing average or a same-day-last-week value are the baseline that a real forecasting method is measured against. If a system’s output is a trailing average, it is producing the comparison, not the answer.
My POS already gives prep recommendations. Is that enough?
It depends on what it is doing underneath. If the recommendation is a trailing average of recent same-weekdays, it will lag every promotion, ignore weather and local events, and get noticeably less reliable in short time windows and at lower volume. That may be acceptable for stable, high volume items and unacceptable for anything with a tight hold window.
How much history does a forecast need?
More history helps, but a new location is not blocked. Patterns from comparable locations in the same brand can carry a store through its early weeks, which is generally better than a four-week window made up mostly of opening-period behavior.
Does this work for a smaller group?
Yes, and the arithmetic argues it matters more, not less. Averaging degrades as counts get smaller, so lower volume operations get the least reliable numbers from the simplest method while having the least margin to absorb the error.
Do I need clean inventory and invoice data first?
No. Prep forecasting reads sales history and recipe yields. Costing and variance reporting depend on invoices, pack sizes, and unit conversions. Those are different data paths, and a problem in the second one does not have to hold up the first.
Bottom Line
A four-week average is not a bad tool. It is a benchmark being used as an answer. It assumes recent same-weekdays are interchangeable, it cannot explain or anticipate anything, and it degrades exactly where prep decisions are tightest: in short time windows, on small counts, at lower volume locations, and in new stores.
If a kitchen team has quietly gone back to guessing, that is usually not a discipline problem. It is a reasonable response to a number that has not earned their trust.
Turning sales history and recipe yields into a prep number a kitchen can actually rely on, at the granularity the product requires, is the work we do at ClearCOGS. If you want to see what that looks like against your own numbers, Let’s Talk.
Sources
- Rob J Hyndman and George Athanasopoulos, Forecasting: Principles and Practice, 2nd edition, OTexts. Section 12.2, Time series of counts.
