Star-plowhorse-dog-puzzle matrix: the menu before and after you measure it

The star-plowhorse-dog-puzzle matrix sorts every dish into four boxes by crossing contribution margin in currency against relative popularity inside its own family, and its value sits not in the chart but in what the chart forces you to do next: reprice the plowhorses, push the puzzles into the top third of the page, rescue or kill the dogs and shield the stars from any quality cut. A 60-dish menu measured this way typically moves gross margin 3 to 7 points within 90 days WITHOUT a general price rise, because the money never comes from charging more; it comes from selling differently. You need two figures per dish — costed recipe and units sold last quarter — plus a decision rule written down BEFORE anyone looks at the output.
A 140-seat steakhouse in Bogotá sold 3,100 monthly units of a loin that returned 9,800 pesos of margin per plate, and 410 units of an octopus that returned 31,400. The owner swore the octopus was the problem because it would not move. The menu was the problem: the octopus sat on the second-to-last page, set in the same type size as the sides, no photo, and not one word explaining why it cost what it cost.
Kasavana and Smith published this four-box model at Michigan State University in 1982, and it has outlived forty-four years of management fashion because it does exactly one thing well: it separates what sells from what earns. Two axes, four quadrants, no poetic interpretation required.
Now the uncomfortable part. Classic menu engineering ignores station time, labour cost per plate and differential waste, so a dish can land as a star on the grid while it quietly wrecks your line on a Friday at nine. We still run it, with one correction you will find in step 6.
Side-by-side comparison
| Unmeasured menu (before) | Matrix applied (after) | |
|---|---|---|
| Food gross margin | ✕62% average, nobody knows which dish holds it up | ✓66-69% after 90 days, contribution traced per dish |
| Dishes carrying 80% of margin | ✕Unknown; assumed to be the 6 best sellers | ✓11 of 58 identified, protected and featured |
| Food cost of the anchor dish | ✕38% on the top seller, above the ceiling | ✓29-31% after recosting and portion adjustment |
| Menu size | ✕58 dishes, 14 selling under 20 units/month | ✓42 dishes; the 16 dead ones cut or reworked |
| Average food check | ✕48,200 COP, flat for 5 quarters | ✓53,900 COP (+11.8%) with no general price rise |
| Guest decision time | ✕4 min 20 s on average, with questions to the server | ✓2 min 40 s; puzzles get ordered unprompted |
| Inventory shrink | ✕4.9% of food cost, dormant references | ✓2.6%; fewer SKUs and steadier rotation |
Step 1: pull ninety days of sales by dish and by family
The first deliverable is a table with four columns per dish: units sold, menu price, raw food cost, and the family it belongs to. Ninety days, not thirty, because a single month captures one promotion or one rainy week and lies to you. At that 140-seat Bogotá steakhouse, the file came back with 3,100 monthly units of tenderloin and 410 of octopus, and that gap alone was enough for the owner to blame the octopus before ever looking at margin. You verify the deliverable by adding up sales across every dish: if the total doesn't match the register report for the same period within 2%, you have dishes off the list or combos your POS refuses to break apart. Fix that before moving on, because the entire matrix rests on that sum. The money axis is margin in CURRENCY, and half the menus I review get lost right here.
Step 2: calculate contribution margin in currency, not in percentage
Menu price minus raw food cost, dish by dish, with no payroll, rent, or utilities loaded onto the plate: those belong to break-even, not to the dish. The tenderloin left 9,800 pesos per unit and the octopus 31,400, meaning the octopus tripled unit contribution while the owner kept calling it the weak one. Now multiply: 3,100 times 9,800 gives 30.4 million a month against 12.9 million from the octopus. Both figures matter and they say different things. A dish carrying 72% margin that sells 40 units puts less cash in the drawer than one at 58% selling 900, and the percentage alone will never tell you that. Done when every dish has unit contribution and total contribution. Kasavana and Smith did not use the simple average and neither should you. The rule, published in 1982 out of Michigan State University, reads: one divided by the number of dishes in the family, multiplied by 0.70.
Step 3: set the popularity threshold at 70% of the family's average sale
If your entrée family holds 12 dishes, expected average sale is 8.3% and the threshold lands at 5.8%. Any dish above 5.8% of that family's units is POPULAR. The simple average over-punishes long menus: with 12 dishes, demanding 8.3% condemns half of them by arithmetic rather than by performance. And the comparison happens WITHIN the family, never against the whole menu, because crossing desserts with entrées produces four beef stars and a list of dogs that are perfectly ordinary desserts selling like desserts. Deliverable: one column with the threshold per family and a yes/no popularity flag per dish. High margin plus high popularity gives you a STAR: don't touch it, don't discount it, shield it from a supplier change. Low margin with high popularity is a PLOWHORSE, the dish that fills the room and doesn't cover rent.
Step 4: cross the two axes and sort the four quadrants
High margin with low popularity is a PUZZLE, and that's where the octopus sat: 31,400 pesos of contribution buried on the second-to-last page, same type size as the sides, no photo, not one line explaining why it cost what it cost. Low margin and low popularity is a DOG. Four boxes, zero poetic interpretation. The matrix earns nothing from the drawing; it earns everything from what it forces you to do the following Monday. Verify that the dish count across the four quadrants equals the full menu and that no family landed entirely in one box, which signals a badly calculated threshold. You raise the plowhorse's price, full stop. It's the dish people already order, with proven demand and less price sensitivity than you fear; start at 6% to 8% and measure four weeks of units before moving further. The puzzle gets repositioned instead: first third of the menu, its own type size, a description that justifies the price through origin or technique, and a photo.
Step 5: raise plowhorse prices and push puzzles into the first third
Cornell measured up to 30% more sales for a dish carrying an image, plus roughly 6.5% per dish when the photography is professional. At that steakhouse, moving the octopus to page one with a photo and dropping the tenderloin one line would have shifted the mix without touching the kitchen. Oracle NetSuite puts sustained profit improvement from well-executed menu engineering at 10% to 15%, and that is precisely the lever. Done when every plowhorse has a new price and every puzzle a new position. Here's the uncomfortable part classic menu engineering keeps quiet: the matrix ignores station time, labor cost per dish, and differential waste, so a dish can register as a star while sinking your kitchen on Friday at nine. The adjustment is simple and we apply it every time: time the pass on your ten best sellers, calculate cost per minute on your line —with the US restaurant base hourly wage already at 14.20 dollars after climbing 4% in 2024, according to 7shifts, a cook's minute stopped being free— and subtract that cost from margin.
Step 6: adjust for station time, labor, and differential waste
Add real waste, not theoretical waste. A dish that occupies the flattop 11 minutes at peak blocks two tickets: its true contribution is its own minus what never left the pass. With that column, two or three stars turn back into plowhorses. Four failures explain almost every useless matrix that reaches my desk. First: using food cost percentage as the money axis, which sorts the menu backwards and rewards dishes that move no cash. Second: mixing families, with the dessert-shaped dogs already described. Third: killing the dog on reflex. The dog doesn't always die; when it carries a brand promise —the only vegan option, grandfather's dish that brings in a table of twelve— it becomes acquisition cost and it stays, though with a cheaper recipe. Fourth, and the most expensive: running the matrix once and filing it. Input prices move; fresh salmon fell 3% and frozen shrimp 6.6% in March 2024 according to SeafoodSource, and a swing like that reorders quadrants on its own.
Common mistakes that ruin the matrix before you apply it
Recalculate quarterly, by calendar, not by mood. Check six points and you'll know the work is finished. One: your table's sales total matches the register within 2%. Two: every dish carries unit contribution in currency and total contribution. Three: every family has its own threshold built with the 70% rule, not with the average. Four: every plowhorse has its new price applied and the application date written down. Five: every puzzle changed position, description, or photo, and you can point to which of the three. Six: an adjustment column for pass time and waste exists on at least your ten best sellers. If all six hold, measure total dining-room contribution at sixty days against the baseline; the reasonable range we work with at Masterestaurant, and the one Diego F. Parra defends in every implementation, is that 10% to 15% improvement Oracle NetSuite documents. Put the next recalculation date on the calendar today.
What actually changes once you measure?
The money axis is contribution margin in CURRENCY, never food-cost percentage. A dish at 72% margin selling 40 units banks less cash than one at 58% selling 900, and the percentage alone will never tell you that.
Popularity is measured INSIDE the family, not against the whole menu. Comparing a dessert with a main course manufactures four meat stars and a dog list made of perfectly ordinary desserts. The popularity threshold is not the average: it is 70% of the family's expected mean share (1 divided by the number of dishes, times 0.70). That is the Kasavana and Smith rule, and the plain average over-punishes long menus. A dog does not always die. When it carries a brand promise — the only vegan option, grandmother's dish that brings the whole family — it becomes a positioning cost, decided with a cold head rather than with nostalgia. Plowhorse price rises go in steps of 4 to 7%, never in one jump, and you read real elasticity in the POS three weeks later.
What actually changes once you measure — in practice
Physical redesign matters as much as the arithmetic: order and visual weight shift sales mix by 8 to 15% without touching a single price.
Instinct versus matrix: where each one wins
Before: a menu defended by instinctStarting point
- Prices set by looking at the place next door, never by recipe costing
- No unit sales data beyond the POS top ten
- Inherited dishes nobody dares remove because regulars still ask
- Menu ordered by product category, with no visual or economic hierarchy
- Blended food cost of 38%, never broken down by family
After: the menu as financial structureMasterestaurant
- Every dish carries a contribution margin in currency, not a percentage
- Last quarter's sales mix crossed against that margin
- Decision rule written and signed before results are visible
- Layout that pushes puzzles and shields stars in the hot zone
- Quarterly review with a numeric checkpoint per quadrant
Side-by-side comparison
| Unmeasured menu (before) | Matrix applied (after) | |
|---|---|---|
| Food gross margin | ✕62% average, nobody knows which dish holds it up | ✓66-69% after 90 days, contribution traced per dish |
| Dishes carrying 80% of margin | ✕Unknown; assumed to be the 6 best sellers | ✓11 of 58 identified, protected and featured |
| Food cost of the anchor dish | ✕38% on the top seller, above the ceiling | ✓29-31% after recosting and portion adjustment |
| Menu size | ✕58 dishes, 14 selling under 20 units/month | ✓42 dishes; the 16 dead ones cut or reworked |
| Average food check | ✕48,200 COP, flat for 5 quarters | ✓53,900 COP (+11.8%) with no general price rise |
| Guest decision time | ✕4 min 20 s on average, with questions to the server | ✓2 min 40 s; puzzles get ordered unprompted |
| Inventory shrink | ✕4.9% of food cost, dormant references | ✓2.6%; fewer SKUs and steadier rotation |
The numbers behind the decision
“We walked in convinced we had to raise everything 10% because beef had run away from us. Diego made us cost all 58 dishes and cross them against the quarter's sales before touching anything. Eleven dishes turned out to generate nearly all the margin, and fourteen never reached twenty sales a month. We raised prices on six, moved four into the top third of the page and cut sixteen. Food check went from 48,200 to 53,900 pesos in one quarter and gross margin gained five points. Nobody complained about price, because we never touched what people compare.”
Building the matrix step by step, with a deliverable and a checkpoint each
Before the first number you need a costed recipe for every dish with grammages actually weighed in the kitchen, a 90-day unit sales report per dish exported from the POS, and the current price list with tax split out. DELIVERABLE: one sheet, one row per dish, five columns — name, family, net price, recipe cost, quarterly units. CHECKPOINT: the sum of net price times units must reconcile with food revenue on the P&L within 2%. If it does not, you have off-system sales, unlogged comps or badly exploded combos, and any matrix built on that data will lie to you with two decimal places. Common error: using a recipe cost from two years ago, when the supplier has changed three times since.
Contribution margin equals net price minus recipe cost. In currency, not as a percentage, and with no payroll, rent or utilities loaded onto the plate: those belong in the break-even calculation, never in the recipe cost. DELIVERABLE: a unit margin column and a quarterly total margin column. CHECKPOINT: sort by total margin descending and mark where 80% accumulates; on a healthy menu that happens between dish 8 and dish 15. Needing 30 dishes to reach 80% means the menu is diluted and most of your work will be pruning. Common error: leaving tax inside the price while the cost sits net of tax, which inflates every margin by 8 to 19 points depending on the country.
The money threshold is the family's weighted average contribution margin. The popularity threshold is the expected mean share per dish in that family multiplied by 0.70: ten dishes means a 10% expected share, so the line falls at 7%. Write both numbers on the sheet and date them. DELIVERABLE: two documented constants per family. CHECKPOINT: both thresholds must sit in a saved file before you classify the first dish. Common error: nudging the threshold down when the chef's favourite lands in the dog box. That retroactive tweak is the most elegant way to change nothing while believing you ran menu engineering.
High margin and high popularity, star. Low margin and high popularity, plowhorse. High margin and low popularity, puzzle. Low margin and low popularity, dog. DELIVERABLE: the full menu tagged, with a count per quadrant and per family. CHECKPOINT: a balanced menu lands near 20% stars, 30% plowhorses, 25% puzzles and 25% dogs; if dogs exceed 35% of the list, your issue is menu length rather than pricing. Common error: classifying against the whole menu instead of by family, which turns starters and desserts into dogs by mathematical construction rather than by real performance.
Stars: leave the recipe and the portion alone, and raise price at most 3% a year when cost pushes. Plowhorses: lift 4 to 7% or recost the portion until food cost drops under 32%. Puzzles: move them into the top third of their section, give them a description naming the producer or origin, and drill the pitch with the floor team. Dogs: out, unless they hold a brand promise. DELIVERABLE: an action list with an owner and a date. CHECKPOINT: three weeks later the promoted puzzle must gain at least 25% in units; flat units mean the pitch failed, not the dish. Common error: repricing and redesigning on the same day, which destroys your reading of which lever worked.
This is where the classic grid falls short, and where I got it wrong for years by pushing puzzles the line could not sustain at peak. Add a third column: direct labour minutes per plate, timed with a stopwatch during real service rather than estimated. Divide contribution margin by those minutes and you get margin per line minute. DELIVERABLE: a margin-per-minute ranking. CHECKPOINT: no promoted dish may sit below the 40th percentile of that list. A glorious puzzle that eats eleven minutes of griddle during a 200-cover service is no puzzle at all: it is a bottleneck with good photography, and it will cost you more in delayed tickets than it returns in margin.
Arithmetic changes the decision; layout changes behaviour. Put puzzles in the top third of each section and inside the box when there is a box, leave stars where they are because they already work, drop the currency symbol, run prices at the end of the description instead of in a column, and kill the price ladder down the right margin. Menu price psychology here is cheap arithmetic with real effect. DELIVERABLE: a new menu, printed and digital, with an effective date. CHECKPOINT: measure sales mix at 21 days; the shift toward puzzles and stars should land between 8 and 15%. Common error: redesigning without measuring, so the next cost crisis sends you straight back to instinct.
Rerun the calculation every 90 days in the same format and archive each run. The series beats any single run because it shows movement: a puzzle that fails to become a star across two consecutive quarters is not a badly communicated puzzle, it is a dish your guest does not want. DELIVERABLE: a history with food gross margin, average check and quadrant counts, quarter by quarter. CHECKPOINT: gross margin should gain 3 to 7 accumulated points across the first two cycles; six months of flat margin with the matrix running points at purchasing or portioning, not at the menu, and that calls for a different tool.
And with AI?
Optimize menu engineering, descriptions and the photos that sell most. Diego F. Parra is an expert in AI applied to restaurants.
Free tools to apply this now
Masterestaurant tools for this guide
You can build the matrix in a spreadsheet, and starting there is genuinely useful because it teaches the mechanics before you automate them. Past forty dishes, or once a group runs more than one location, the spreadsheet starts failing on version control and the work turns into file archaeology.
These three pieces of the Masterestaurant method cover the three moments of the decision: the business model that justifies the menu, the projection of what the new margin does to profit, and cash control while the transition plays out.
Frequently asked questions about the star-plowhorse-dog-puzzle matrix
How often should I rerun the star-plowhorse-dog-puzzle matrix?
How often should I rerun the star-plowhorse-dog-puzzle matrix?
Every 90 days as a rule, and immediately whenever a core input moves more than 15% or the menu changes. Less than a quarter leaves no time for the sales mix to settle after a price or layout change, and more than six months means deciding on data that already expired.
What do I do with a dog dish that regulars always order?
What do I do with a dog dish that regulars always order?
Put a number on the sentiment. Work out how much margin you forfeit per year to keep it, then decide whether that positioning cost buys enough loyalty. If it does, keep it off the printed menu and alive as a server suggestion; it stops occupying expensive space and stays available to whoever asks.
Does the matrix work when my menu rotates weekly by season?
Does the matrix work when my menu rotates weekly by season?
It works by family rather than by individual dish. Classify the category — seafood starters, grilled mains, baked desserts — and apply the rule to the archetype instead of the specific recipe. On seasonal menus the useful figure is the family's weighted average margin and its rotation, which stays comparable across cycles.
Is the matrix the same thing as a 32% food cost analysis?
Is the matrix the same thing as a 32% food cost analysis?
No. Food cost per dish is a control ceiling — 32% maximum, and 32% is already high — while the matrix decides which dishes to push based on the cash they return and how often they sell. A dish at 30% food cost can still be a dog when nobody orders it, and the percentage alone will never show you that.
Sector data 2026 (official sources)
Verifiable industry benchmarks from official, non-commercial sources (government, industry associations, market research) - not competitors.
| Metric | Benchmark 2026 | Source |
|---|---|---|
| Personas con alergias alimentarias comprobadas (EE. UU.) | Más de 30 millones | US FDA / FARE — 2024 |
| Visitas anuales a urgencias por alergias alimentarias (EE. UU.) | Más de 200.000 al año | Food Allergy Research & Education (FARE) |
| Consumidores que evitan productos con alérgenos mayores (EE. UU.) | 25% de los consumidores | Food Allergy Research & Education (FARE) |
| Lealtad de comensales con alergias alimentarias | 36% siempre visita el mismo lugar vs 17% sin alergias | Estudio Food Allergy and Foodservice — PMC |
| Umbral de la regla de etiquetado de calorías en el menú (FDA) | Cadenas con 20 o más locales | US Food and Drug Administration — Menu Labeling |
| Reducción de calorías por el etiquetado en el menú | ≈7,3% menos de calorías | US FDA / estudios de menu labeling |
Related content
Grow your restaurant with the Masterestaurant method
Applied in +8.400 restaurants across 43 countries.
