What Are Diners Really Saying? Using AI to Analyse Your Reviews

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for What Are Diners Really Saying? Using AI to Analyse Your Reviews.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for What Are Diners Really Saying? Using AI to Analyse Your Reviews.

Collect six to twelve months of reviews from every platform into one sheet. Give the AI a fixed list of themes (each dish, service, waiting time, value, atmosphere, cleanliness, booking), have it tag every review with a short quote as evidence, then count the tags by month in a spreadsheet. Check 20 tags yourself before trusting any total.

The quote-as-evidence step is what makes this trustworthy. Ask an assistant "what are people saying about us?" and you'll get a confident paragraph you can't check. Ask it to tag each review against your list and show the words that justify each tag, and you get something you can count, audit and act on.

Follow me on Instagram@sagnikteaches

Gathering every review into one sheet

Most restaurants have reviews scattered across Google, a booking platform, one or two review sites, delivery apps and their own feedback emails. Get them into one sheet with these columns: review ID, date, platform, star rating, review text. Leave the reviewer's name out; you don't need it for analysis, and there's no reason to paste it into an AI tool.

Connect on LinkedInSagnik Bhattacharya
  • Google. Owners and managers can request a Business Profile export through Google Takeout. It arrives as a compressed folder with review data in JSON, a format built for software rather than people. Excel opens it through Data, Get Data, From File, From JSON, then expanding the records into columns. If that feels like too much, copying recent reviews by hand is fine for a few dozen.
  • Booking platforms and review sites. Check each dashboard for an export. Where there isn't one, copy and paste; it's tedious but a one-off.
  • Delivery apps. Ratings and comments often sit in the merchant dashboard. Keep them in a separate platform column, because delivery complaints (late, cold, crushed) are a different problem from dining-room ones.
  • Your own feedback. Emails, comment cards and post-visit surveys belong here too. They're often more detailed than public reviews.

If the result is messy, with dates in three formats and duplicate rows, cleaning messy data covers the fixes before you start.

Subscribe on YouTube@codingliquids

Deciding which reviews to include

More isn't always better. Set the window before you start:

  • Six to twelve months is usually right. Shorter gives too few reviews to see patterns; longer mixes in a restaurant that may no longer exist.
  • Start after any big change. A new head chef, a refit or a new menu resets what diners are reacting to. Analysing across the change blurs both halves, unless comparing before and after is the point.
  • Keep reviews with no text out of the theme analysis, but count them separately. A run of text-free one-star ratings is a signal in itself.
  • Mark private feedback so you can see whether it tells a different story from public reviews. It often does: people complain about value more readily in private.

A theme list built for a restaurant

The single most important decision is the list of themes, sometimes called a codebook. Fix it before tagging, so the AI sorts reviews into your categories rather than inventing new ones each batch. A starting point:

ThemeWhat countsExample phrase
Food: [dish name]Any comment about a named dish; one tag per dish mentioned"the short rib was incredible"
Food: generalFood comments with no dish named"food was a bit bland"
ServiceFriendliness, attentiveness, knowledge"our server couldn't do enough"
SpeedWaits for a table, food, drinks or the bill"took 20 minutes to get the bill"
ValuePrice against what was received"a lot for what you get"
AtmosphereNoise, music, lighting, room temperature, decor"too loud to talk"
CleanlinessToilets, tables, cutlery, visible kitchen"toilets need attention"
Booking and arrivalLost bookings, table not ready, phone manner"booking had disappeared"
Dietary and allergensHow requests were handled"they were brilliant with my coeliac daughter"
Delivery and takeawayPackaging, lateness, temperature on arrival"arrived lukewarm"

Each tag also gets a sentiment: positive, negative or mixed. Add your menu's dish names to the prompt so "the beef" and "short rib" land on the same dish.

A takeaway, or a restaurant doing most of its trade through delivery apps, needs the delivery row split up, because "delivery" on its own hides four different fixes: missing or wrong items (packing), temperature (packaging and wait times), lateness (often the courier, but worth counting), and presentation on arrival (crushed boxes, leaking sauces). In an illustrative quarter of 90 app reviews at a pizza takeaway, missing items came up 14 times, and 11 of those were orders where a dip or a side had been added as an extra. That points at the packing checklist, not the kitchen, and a single "delivery" theme would never have shown it.

The tagging prompt

Work in batches of 40 to 60 reviews. Larger batches make tags sloppier towards the end of the batch.

You are tagging restaurant reviews for analysis.
Themes (use ONLY these): [paste codebook themes]
Our dishes: [list, with common nicknames: "beef" = Short rib]

For each review, output one row per theme mentioned:
Review ID | Theme | Sentiment (Positive/Negative/Mixed) | Evidence
Evidence = the exact words from the review, under 15 words.
If a review mentions no theme, output: Review ID | None | - | -
Do not summarise. Do not count. Do not add themes.
Sarcasm ("lovely 50-minute wait") is Negative.

[paste batch: Review ID | Date | Platform | Stars | Text]

Here's what a batch returns for three reviews (sample output, illustrative). The reviews: R017, "Beef was gorgeous as always. Shame we waited 25 minutes for the bill on a quiet Tuesday." R018, four stars, "Nice enough. Bit pricey for a burger that was dry." R019, "Lovely 50-minute wait for our mains, really made the evening."

R017 | Food: Short rib | Positive | "Beef was gorgeous as always"
R017 | Speed          | Negative | "waited 25 minutes for the bill"
R018 | Value          | Negative | "Bit pricey"
R018 | Food: general  | Negative | "a burger that was dry"
R019 | Speed          | Positive | "Lovely 50-minute wait"

R017 is right, including the nickname "beef". R018's second row should be tagged to the burger, since a dish is named. The model missed it because the dish list gave the menu name, "Grill stack", but not the word diners actually use, so add "burger" = Grill stack. R019 is the sarcasm case, tagged positive despite the instruction. Rather than argue with the model, add an example to the prompt: "'Lovely 50-minute wait' = Speed, Negative". A worked example in the prompt corrects this kind of error more reliably than a rule does.

"Do not count" is deliberate. Assistants are unreliable at counting across long lists, so paste the tagged rows into your sheet and let a pivot table or COUNTIFS do the totals. If you'd rather stay inside the assistant, upload the tagged file and ask it to count using code, which is far more reliable than counting in its head.

Checking the tags before you trust the totals

Pick 20 tagged reviews at random and check each tag yourself. Note how many you agree with. In a typical first run you might agree with 17 of 20, and the misses tend to fall into a few patterns:

  • Sarcasm and understatement tagged as positive.
  • Mixed reviews ("great food, shame about the wait") tagged only on the first theme.
  • Value confused with portion size ("huge portions!" tagged as value).
  • Star rating and text disagreeing. A five-star review that says "service was slow but we didn't mind" is still a speed mention.

A few rows from a filled-in check sheet (illustrative):

ReviewAI tagYour tagAgree?Pattern
R042Value, Positive ("huge portions")Food: general, PositiveNoPortion read as value
R051Food: Short rib, PositiveFood: Short rib, Positive; Speed, NegativeNoMixed review, second theme missed
R063Atmosphere, Negative ("couldn't hear each other")SameYes–
R077Service, PositiveSameYes–
R088Value, Positive ("plates piled high")Food: general, PositiveNoPortion read as value

Two of the three misses share a pattern, so the fix goes in the codebook, not the prompt. The Value row changes from "Price against what was received" to "Price against what was received. Portion size on its own is Food, not Value; tag Value only when price or cost is mentioned."

Fix the codebook or prompt for whatever pattern you find, re-run that batch, and check another 20. When you agree with about 18 in 20 or better, the counts are good enough to act on. Keep the check sheet: it's your evidence if someone on the team questions a finding.

Worked example: nine months at a neighbourhood restaurant

Say a neighbourhood restaurant (an illustration) collects 240 reviews from January to September: 100 from January to April and 140 from May to September, after a price rise in early May. Tagged and counted:

ThemeMentionsPositiveNegativeWhat stands behind it
Food (all)16814127Short rib praised 31 times; burger called dry or overcooked 11 times
Service1219823Negatives spread thinly; no single cause
Speed5264638 of the 46 negatives are Friday or Saturday dinners
Value4420245 negatives January-April; 19 May-September
Atmosphere372215Noise, mostly Saturday nights
Cleanliness945All five negatives about the toilets

Two adjustments make the table honest. First, compare rates, not raw counts, because review numbers vary month to month. Value complaints went from 5 per 100 reviews before the price rise to about 14 per 100 after (19 of 140). That's a real shift, not just more reviews. Second, look at where the negatives cluster: speed complaints aren't a general service problem, they're a weekend-dinner capacity problem, which points to the rota and the kitchen rather than to training.

The burger line is worth setting beside your sales data. If it's also a popular, low-margin dish, menu engineering with AI will flag it as a Ploughhorse, and eleven complaints suggest fixing it before it drags on the rest of the menu.

From findings to three changes

Resist fixing everything. Pick the three findings with the most negative mentions per 100 reviews that you can actually change, and write each as an action with a measure:

Finding: Speed complaints concentrated on Fri/Sat dinner (38 of 46)
Change:  Add a cook 7-9pm Fri/Sat; tell tables when kitchen runs behind
Measure: Speed negatives per 100 weekend reviews, next 8 weeks
Owner:   [name]    Review date: [date]

Finding: Value negatives up from 5 to ~14 per 100 since May prices
Change:  [e.g. restore the side with mains, or revisit two prices]
Measure: Value negatives per 100 reviews, next 8 weeks
Owner:   [name]    Review date: [date]

Then re-run the tagging on new reviews every month or quarter using the same codebook, so the numbers stay comparable. Changing the themes midway breaks the trend line.

Checking the first change eight weeks on, with illustrative numbers: before the extra weekend cook, 38 speed negatives came from about 120 weekend-dinner reviews, roughly 32 per 100. In the eight weeks after, 30 weekend-dinner reviews contained 4 speed negatives, about 13 per 100. That's a clear move in the right direction, but 30 reviews is a small base, so keep the change and look again at the next quarter before calling it fixed. If the rate had stayed near 30 per 100, the problem would lie elsewhere: the pass, the order in which tickets are fired, or the bar, not the number of cooks.

The positive tags are useful too. The exact words diners use about your best dishes are good raw material for your menu and posts; turning customer reviews into marketing copy shows how to use them without misquoting anyone. And the recurring complaints feed straight into how you reply: answering restaurant complaints with AI covers the replies themselves.

Sharing the results with the team

How you present the findings decides whether anything changes. A few habits help:

  • One page, not a spreadsheet. The three findings, the numbers behind each, two or three real quotes, and the actions. Kitchen and floor teams read a page; they don't open a pivot table.
  • Lead with what diners love. Thirty-one mentions of the short rib is news the kitchen should hear first, and it makes the harder findings easier to take.
  • Keep named staff out of it. Reviews that mention a server by name, good or bad, are for a private conversation with that person, not the team sheet.
  • Show the trend next time. The second report is more powerful than the first, because the team can see whether their changes moved the numbers.

The one-page sheet for the neighbourhood restaurant, filled in (illustrative):

What diners love: the short rib (31 mentions: "the best thing on the menu", "worth the trip on its own"); friendly service (98 positive mentions). Three things we're changing: 1. Weekend waits: 38 of 46 speed complaints are Friday and Saturday dinners, so there's an extra cook 7-9pm from next week. 2. Value since the May prices: complaints are up from 5 to about 14 per 100 reviews, so we're putting the side back with mains. 3. The burger: 11 reviews call it dry or overcooked, so the head chef is retesting the cook time and patty size. Next check: the second week of December.

It takes about ten minutes to read aloud at a staff meeting, and every number on it can be traced back to the sheet.

Time-wise, expect three or four hours for the first run on a couple of hundred reviews, most of it gathering and checking. After that, a monthly update on new reviews takes about half an hour.

What reviews can't tell you

  • Reviewers aren't typical diners. People write reviews after unusually good or unusually bad visits. A theme's share of reviews isn't its share of all visits.
  • Small numbers mislead. Five toilet complaints in nine months is worth fixing because it's cheap to fix, not because it's statistically meaningful.
  • One loud review isn't a trend. The long, furious review everyone on the team remembers may be the only one of its kind. The counts keep it in proportion.
  • Silence isn't satisfaction. If nobody mentions your desserts, that might mean they're fine or that few people order them. Check sales before concluding anything. Say desserts get three mentions in 240 reviews and the till shows only one table in twelve ordering one: that's a menu and upselling question worth asking, not evidence that the desserts are fine.

For feedback you collect yourself through post-visit surveys, the same method applies with a few differences in question design; analysing customer feedback surveys with AI covers those. If you'd rather not build the collection yourself, AI review management tools pull reviews from several platforms into one dashboard, though you'll still want your own codebook for analysis that reflects your menu.

Further reads

Sources: Google Business Profile and Google Takeout export information; Microsoft Excel JSON import (Get Data). Worked example figures are illustrative. Checked September 2026.

Want your reviews turned into a to-do list?

On a 1:1 call we'll gather your reviews from each platform, build a theme list that fits your menu and service, and set up a monthly count your team can run without help.

Book a 1:1 call with me