Collect six to twelve months of reviews from every platform into one sheet. Give the AI a fixed list of themes (each dish, service, waiting time, value, atmosphere, cleanliness, booking), have it tag every review with a short quote as evidence, then count the tags by month in a spreadsheet. Check 20 tags yourself before trusting any total.
The quote-as-evidence step is what makes this trustworthy. Ask an assistant "what are people saying about us?" and you'll get a confident paragraph you can't check. Ask it to tag each review against your list and show the words that justify each tag, and you get something you can count, audit and act on.
Gathering every review into one sheet
Most restaurants have reviews scattered across Google, a booking platform, one or two review sites, delivery apps and their own feedback emails. Get them into one sheet with these columns: review ID, date, platform, star rating, review text. Leave the reviewer's name out; you don't need it for analysis, and there's no reason to paste it into an AI tool.
- Google. Owners and managers can request a Business Profile export through Google Takeout. It arrives as a compressed folder with review data in JSON, a format built for software rather than people. Excel opens it through Data, Get Data, From File, From JSON, then expanding the records into columns. If that feels like too much, copying recent reviews by hand is fine for a few dozen.
- Booking platforms and review sites. Check each dashboard for an export. Where there isn't one, copy and paste; it's tedious but a one-off.
- Delivery apps. Ratings and comments often sit in the merchant dashboard. Keep them in a separate platform column, because delivery complaints (late, cold, crushed) are a different problem from dining-room ones.
- Your own feedback. Emails, comment cards and post-visit surveys belong here too. They're often more detailed than public reviews.
If the result is messy, with dates in three formats and duplicate rows, cleaning messy data covers the fixes before you start.
Deciding which reviews to include
More isn't always better. Set the window before you start:
- Six to twelve months is usually right. Shorter gives too few reviews to see patterns; longer mixes in a restaurant that may no longer exist.
- Start after any big change. A new head chef, a refit or a new menu resets what diners are reacting to. Analysing across the change blurs both halves, unless comparing before and after is the point.
- Keep reviews with no text out of the theme analysis, but count them separately. A run of text-free one-star ratings is a signal in itself.
- Mark private feedback so you can see whether it tells a different story from public reviews. It often does: people complain about value more readily in private.
A theme list built for a restaurant
The single most important decision is the list of themes, sometimes called a codebook. Fix it before tagging, so the AI sorts reviews into your categories rather than inventing new ones each batch. A starting point:
| Theme | What counts | Example phrase |
|---|---|---|
| Food: [dish name] | Any comment about a named dish; one tag per dish mentioned | "the short rib was incredible" |
| Food: general | Food comments with no dish named | "food was a bit bland" |
| Service | Friendliness, attentiveness, knowledge | "our server couldn't do enough" |
| Speed | Waits for a table, food, drinks or the bill | "took 20 minutes to get the bill" |
| Value | Price against what was received | "a lot for what you get" |
| Atmosphere | Noise, music, lighting, room temperature, decor | "too loud to talk" |
| Cleanliness | Toilets, tables, cutlery, visible kitchen | "toilets need attention" |
| Booking and arrival | Lost bookings, table not ready, phone manner | "booking had disappeared" |
| Dietary and allergens | How requests were handled | "they were brilliant with my coeliac daughter" |
| Delivery and takeaway | Packaging, lateness, temperature on arrival | "arrived lukewarm" |
Each tag also gets a sentiment: positive, negative or mixed. Add your menu's dish names to the prompt so "the beef" and "short rib" land on the same dish.
A takeaway, or a restaurant doing most of its trade through delivery apps, needs the delivery row split up, because "delivery" on its own hides four different fixes: missing or wrong items (packing), temperature (packaging and wait times), lateness (often the courier, but worth counting), and presentation on arrival (crushed boxes, leaking sauces). In an illustrative quarter of 90 app reviews at a pizza takeaway, missing items came up 14 times, and 11 of those were orders where a dip or a side had been added as an extra. That points at the packing checklist, not the kitchen, and a single "delivery" theme would never have shown it.
The tagging prompt
Work in batches of 40 to 60 reviews. Larger batches make tags sloppier towards the end of the batch.
You are tagging restaurant reviews for analysis.
Themes (use ONLY these): [paste codebook themes]
Our dishes: [list, with common nicknames: "beef" = Short rib]
For each review, output one row per theme mentioned:
Review ID | Theme | Sentiment (Positive/Negative/Mixed) | Evidence
Evidence = the exact words from the review, under 15 words.
If a review mentions no theme, output: Review ID | None | - | -
Do not summarise. Do not count. Do not add themes.
Sarcasm ("lovely 50-minute wait") is Negative.
[paste batch: Review ID | Date | Platform | Stars | Text]
Here's what a batch returns for three reviews (sample output, illustrative). The reviews: R017, "Beef was gorgeous as always. Shame we waited 25 minutes for the bill on a quiet Tuesday." R018, four stars, "Nice enough. Bit pricey for a burger that was dry." R019, "Lovely 50-minute wait for our mains, really made the evening."
R017 | Food: Short rib | Positive | "Beef was gorgeous as always"
R017 | Speed | Negative | "waited 25 minutes for the bill"
R018 | Value | Negative | "Bit pricey"
R018 | Food: general | Negative | "a burger that was dry"
R019 | Speed | Positive | "Lovely 50-minute wait"
R017 is right, including the nickname "beef". R018's second row should be tagged to the burger, since a dish is named. The model missed it because the dish list gave the menu name, "Grill stack", but not the word diners actually use, so add "burger" = Grill stack. R019 is the sarcasm case, tagged positive despite the instruction. Rather than argue with the model, add an example to the prompt: "'Lovely 50-minute wait' = Speed, Negative". A worked example in the prompt corrects this kind of error more reliably than a rule does.
"Do not count" is deliberate. Assistants are unreliable at counting across long lists, so paste the tagged rows into your sheet and let a pivot table or COUNTIFS do the totals. If you'd rather stay inside the assistant, upload the tagged file and ask it to count using code, which is far more reliable than counting in its head.
Checking the tags before you trust the totals
Pick 20 tagged reviews at random and check each tag yourself. Note how many you agree with. In a typical first run you might agree with 17 of 20, and the misses tend to fall into a few patterns:
- Sarcasm and understatement tagged as positive.
- Mixed reviews ("great food, shame about the wait") tagged only on the first theme.
- Value confused with portion size ("huge portions!" tagged as value).
- Star rating and text disagreeing. A five-star review that says "service was slow but we didn't mind" is still a speed mention.
A few rows from a filled-in check sheet (illustrative):
| Review | AI tag | Your tag | Agree? | Pattern |
|---|---|---|---|---|
| R042 | Value, Positive ("huge portions") | Food: general, Positive | No | Portion read as value |
| R051 | Food: Short rib, Positive | Food: Short rib, Positive; Speed, Negative | No | Mixed review, second theme missed |
| R063 | Atmosphere, Negative ("couldn't hear each other") | Same | Yes | – |
| R077 | Service, Positive | Same | Yes | – |
| R088 | Value, Positive ("plates piled high") | Food: general, Positive | No | Portion read as value |
Two of the three misses share a pattern, so the fix goes in the codebook, not the prompt. The Value row changes from "Price against what was received" to "Price against what was received. Portion size on its own is Food, not Value; tag Value only when price or cost is mentioned."
Fix the codebook or prompt for whatever pattern you find, re-run that batch, and check another 20. When you agree with about 18 in 20 or better, the counts are good enough to act on. Keep the check sheet: it's your evidence if someone on the team questions a finding.
Worked example: nine months at a neighbourhood restaurant
Say a neighbourhood restaurant (an illustration) collects 240 reviews from January to September: 100 from January to April and 140 from May to September, after a price rise in early May. Tagged and counted:
| Theme | Mentions | Positive | Negative | What stands behind it |
|---|---|---|---|---|
| Food (all) | 168 | 141 | 27 | Short rib praised 31 times; burger called dry or overcooked 11 times |
| Service | 121 | 98 | 23 | Negatives spread thinly; no single cause |
| Speed | 52 | 6 | 46 | 38 of the 46 negatives are Friday or Saturday dinners |
| Value | 44 | 20 | 24 | 5 negatives January-April; 19 May-September |
| Atmosphere | 37 | 22 | 15 | Noise, mostly Saturday nights |
| Cleanliness | 9 | 4 | 5 | All five negatives about the toilets |
Two adjustments make the table honest. First, compare rates, not raw counts, because review numbers vary month to month. Value complaints went from 5 per 100 reviews before the price rise to about 14 per 100 after (19 of 140). That's a real shift, not just more reviews. Second, look at where the negatives cluster: speed complaints aren't a general service problem, they're a weekend-dinner capacity problem, which points to the rota and the kitchen rather than to training.
The burger line is worth setting beside your sales data. If it's also a popular, low-margin dish, menu engineering with AI will flag it as a Ploughhorse, and eleven complaints suggest fixing it before it drags on the rest of the menu.
From findings to three changes
Resist fixing everything. Pick the three findings with the most negative mentions per 100 reviews that you can actually change, and write each as an action with a measure:
Finding: Speed complaints concentrated on Fri/Sat dinner (38 of 46)
Change: Add a cook 7-9pm Fri/Sat; tell tables when kitchen runs behind
Measure: Speed negatives per 100 weekend reviews, next 8 weeks
Owner: [name] Review date: [date]
Finding: Value negatives up from 5 to ~14 per 100 since May prices
Change: [e.g. restore the side with mains, or revisit two prices]
Measure: Value negatives per 100 reviews, next 8 weeks
Owner: [name] Review date: [date]
Then re-run the tagging on new reviews every month or quarter using the same codebook, so the numbers stay comparable. Changing the themes midway breaks the trend line.
Checking the first change eight weeks on, with illustrative numbers: before the extra weekend cook, 38 speed negatives came from about 120 weekend-dinner reviews, roughly 32 per 100. In the eight weeks after, 30 weekend-dinner reviews contained 4 speed negatives, about 13 per 100. That's a clear move in the right direction, but 30 reviews is a small base, so keep the change and look again at the next quarter before calling it fixed. If the rate had stayed near 30 per 100, the problem would lie elsewhere: the pass, the order in which tickets are fired, or the bar, not the number of cooks.
The positive tags are useful too. The exact words diners use about your best dishes are good raw material for your menu and posts; turning customer reviews into marketing copy shows how to use them without misquoting anyone. And the recurring complaints feed straight into how you reply: answering restaurant complaints with AI covers the replies themselves.
Sharing the results with the team
How you present the findings decides whether anything changes. A few habits help:
- One page, not a spreadsheet. The three findings, the numbers behind each, two or three real quotes, and the actions. Kitchen and floor teams read a page; they don't open a pivot table.
- Lead with what diners love. Thirty-one mentions of the short rib is news the kitchen should hear first, and it makes the harder findings easier to take.
- Keep named staff out of it. Reviews that mention a server by name, good or bad, are for a private conversation with that person, not the team sheet.
- Show the trend next time. The second report is more powerful than the first, because the team can see whether their changes moved the numbers.
The one-page sheet for the neighbourhood restaurant, filled in (illustrative):
What diners love: the short rib (31 mentions: "the best thing on the menu", "worth the trip on its own"); friendly service (98 positive mentions). Three things we're changing: 1. Weekend waits: 38 of 46 speed complaints are Friday and Saturday dinners, so there's an extra cook 7-9pm from next week. 2. Value since the May prices: complaints are up from 5 to about 14 per 100 reviews, so we're putting the side back with mains. 3. The burger: 11 reviews call it dry or overcooked, so the head chef is retesting the cook time and patty size. Next check: the second week of December.
It takes about ten minutes to read aloud at a staff meeting, and every number on it can be traced back to the sheet.
Time-wise, expect three or four hours for the first run on a couple of hundred reviews, most of it gathering and checking. After that, a monthly update on new reviews takes about half an hour.
What reviews can't tell you
- Reviewers aren't typical diners. People write reviews after unusually good or unusually bad visits. A theme's share of reviews isn't its share of all visits.
- Small numbers mislead. Five toilet complaints in nine months is worth fixing because it's cheap to fix, not because it's statistically meaningful.
- One loud review isn't a trend. The long, furious review everyone on the team remembers may be the only one of its kind. The counts keep it in proportion.
- Silence isn't satisfaction. If nobody mentions your desserts, that might mean they're fine or that few people order them. Check sales before concluding anything. Say desserts get three mentions in 240 reviews and the till shows only one table in twelve ordering one: that's a menu and upselling question worth asking, not evidence that the desserts are fine.
For feedback you collect yourself through post-visit surveys, the same method applies with a few differences in question design; analysing customer feedback surveys with AI covers those. If you'd rather not build the collection yourself, AI review management tools pull reviews from several platforms into one dashboard, though you'll still want your own codebook for analysis that reflects your menu.
Further reads
- How to Spot Patterns in Complaints, Returns, and Defects With AI — Spotting patterns in complaints beyond public reviews.
- How to Get More Google Reviews With Automated Review Requests — More reviews means more reliable themes.
- How to Reply to Hotel Reviews With AI Without Sounding Canned — The hotel side of hospitality reviews, reply-focused.
- How to Run a Monthly AI Quality Review in 30 Minutes — A 30-minute monthly routine to keep checks honest.
- Your First 30 Days of AI in a Restaurant, Week by Week — Where review analysis fits in a restaurant's first AI month.
- How Much Does AI Cost a Small Restaurant Each Month? — A line-by-line monthly AI budget for a small restaurant, three priced set-ups, and the staff hours that cost as much as the subscriptions.
- Where AI Actually Saves Time in a Small Restaurant — A task-by-task look at a small restaurant's admin week: where AI cuts real minutes, where it only moves them, and where it adds work.
- What AI Can and Cannot Do for an Independent Café — The desk jobs AI does well in a café, the ones it can't touch, and a 30-minute test on your own till data before you pay for anything.
- 7 AI Mistakes Restaurant Owners Make With Bookings and Reviews — Seven ways restaurant AI goes wrong with tables and reviews, each with a real-looking example, what it costs and how to fix it.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Google Business Profile and Google Takeout export information; Microsoft Excel JSON import (Get Data). Worked example figures are illustrative. Checked September 2026.