How to Analyse Customer Feedback Surveys With AI

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Analyse Customer Feedback Surveys With AI.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Analyse Customer Feedback Surveys With AI.

Export the responses to a spreadsheet, remove names and emails, and ask AI to draft a short list of themes from a sample of about 50 comments. Then tag every comment against that fixed list one row at a time, count the tags with a pivot table, and hand-check about one comment in ten before acting on anything.

The usual mistake is pasting 400 answers into a chat and asking what customers think. You get a fluent summary with percentages the model estimated rather than counted, and quotes that are sometimes stitched together from two or three people. Keep the jobs apart: AI reads and labels, the spreadsheet counts, and you decide what the numbers mean.

Follow me on Instagram@sagnikteaches

A craft-kit subscription box and its 412 survey replies

The method is easiest to follow on one business, so the rest of this page uses an illustrative case. A monthly craft-kit subscription box with about 2,100 subscribers emails a short survey after its twelfth box. It gets 412 responses in a week. The survey has four questions:

Connect on LinkedInSagnik Bhattacharya
  1. How likely are you to recommend us to a friend? (0 to 10)
  2. What did you think of this month's box? (free text)
  3. If you could change one thing, what would it be? (free text)
  4. Did you finish the project? (yes, partly, not yet)

The two free-text questions produce 366 usable comments; the rest are blank or say "n/a". Reading all 366 would take the owner most of a day and she would still have no counts at the end. The goal is a table showing which themes come up, how often, and how they relate to the recommend score, so she can decide what to change for box fourteen.

Subscribe on YouTube@codingliquids

Shape the export so the AI can read it without guessing

Every survey tool can export to CSV or straight into Google Sheets or Excel. Before any AI sees the file, rearrange it into one row per response with these columns:

  • response_id: a short code such as R0001. If your tool's IDs are long strings, make your own. Every tag and every quote will point back to this ID.
  • score: the 0 to 10 answer, as a number.
  • comment_box and comment_change: the two free-text answers, untouched.
  • finished: the multiple-choice answer.
  • plan or tenure: anything you already know about the customer that might explain a pattern, such as months subscribed.

Delete the name, email and address columns, and keep the original export somewhere else. Then skim the comments for personal details people typed into the text boxes. In this survey one parent wrote about her son's diagnosis and why the fine motor work suited him; another gave a phone number and asked for a callback. Those details belong in a follow-up email, not in an AI tool. If you have more than a few, the routine in redacting personal data before sharing documents saves time.

One more change pays off later. If people answered two free-text questions, keep them in separate columns rather than joining them. "What would you change" answers are nearly all complaints or wishes, while "what did you think" answers lean positive, and merging them makes every theme look more mixed than it is.

Draft a codebook from 50 comments before tagging anything

A codebook is simply the fixed list of themes you will tag against, each with a code and a one-line definition. Build it from a random sample, not the first 50 rows, because early responders are often your keenest customers. Sort the sheet by a random number column, copy 50 comments, and use a prompt like this:

Below are 50 comments from a customer survey for a monthly craft-kit subscription box.
Propose a codebook of 8 to 12 themes that together cover most of these comments.
For each theme give: a code (T01, T02...), a short name, a one-sentence definition,
and the response IDs from this sample that fit it.
Rules:
- Themes must be about one thing each (not "delivery and packaging").
- Keep complaints and praise about the same topic as separate themes.
- Add T98 Other and T99 Unclear.
- Do not invent themes that no comment supports.

Comments:
R0187 | The instructions for the macrame bit made no sense, gave up halfway
R0033 | Loved it, finished in one evening, my daughter wants the next one
...

An illustrative first draft back from the assistant looked like this:

T01 Project too difficult - customer found the project hard to complete (R0187, R0290...)
T02 Loved the project - positive about the craft itself (R0033, R0145...)
T03 Missing or damaged items - something absent, short or broken on arrival
T04 Delivery timing - box arrived late or on an awkward day
T05 Value - comments about price or what you get for it
T06 More variety - wants different crafts or fewer repeats
T07 Packaging - comments on amount or type of packaging
T08 Pause or skip - wants to skip a month or pause
T09 Suitable for children - comments about using the kit with kids
T98 Other
T99 Unclear

This is a decent start, and it has three problems you should expect in any first draft. T01 mixes two causes: some people found the project hard, others found the instructions unclear, and those need different fixes. T05 hides two opposite complaints, "too expensive" and "not enough yarn to finish". T09 is a topic, not a theme; half those comments were praise and half said the kit was not suitable for under-tens. After editing, the working codebook had twelve themes: instructions unclear, project too difficult, project too easy, loved the project, missing or damaged items, not enough materials, price too high, delivery timing, more variety, packaging waste, pause or skip, and child suitability, plus Other and Unclear. Write a one-line definition for each and a rule for any pair people could confuse.

Tag every comment against the fixed list

Now the AI does the slow part: reading each comment and choosing codes from your list. You have three reasonable ways to run it, depending on the tools you have.

Option 1: a chat assistant with file upload

ChatGPT Plus and Claude Pro both cost about $20 a month and both accept a spreadsheet upload. Send 80 to 100 comments per message rather than all 366 at once; accuracy drifts on long batches and you want to be able to re-run one batch without redoing the lot. Use the same prompt every time:

Tag each survey comment below using ONLY this codebook: [paste codebook with definitions].
A comment can have up to 3 codes. Rules:
- Tag what the customer says, not what you guess they feel.
- If the customer blames the instructions, use T01 (instructions unclear), not T02.
- Sarcasm is negative: "Great, glue again" = negative.
- If nothing fits, use T98 Other and add a 3-word note.
Return CSV only, one line per comment, in the same order:
response_id,codes,sentiment
Sentiment must be one of: positive, negative, mixed, neutral.

Illustrative output for five rows:

R0187,T01,negative
R0033,T04;T12,positive
R0211,T05;T06,negative
R0302,T98 (wants video tutorial),neutral
R0350,T10,positive

Paste each batch's CSV into a new "tags" tab. Check that the row count matches the batch before you move on; a missing row usually means the model merged two short comments.

Option 2: the AI function inside Google Sheets

If your Google Workspace or Google AI plan includes it, the AI function in Google Sheets lets you put the tagging prompt in a formula next to each comment, such as =AI("Using this codebook: ... return only the codes for this comment", C2). Google's help page says only the first 350 selected cells with AI functions generate at a time and that generation limits apply, so run the column in chunks and paste the results as values once they settle, or they may regenerate differently later. For more prompt patterns inside a sheet, see these Google Sheets AI prompts.

Option 3: Copilot in Excel

Microsoft withdrew the COPILOT worksheet function on 14 September 2026, and cells that used it now show a #NAME? error when they recalculate, according to Microsoft's COPILOT function page. The Copilot pane in Excel can still classify a text column and add the results as a new column. Give it the same codebook and rules, and check a sample exactly as you would with the other options.

Count in the spreadsheet, never in the chat

Once every comment has codes, split the codes into one row per code (a comment with two codes becomes two rows) and build a pivot table: themes down the side, count of response IDs, and the average recommend score of the people who mentioned each theme. The illustrative result for the craft box:

ThemeCommentsShare of 366 commentersAverage recommend score
Loved the project11832%8.9
Instructions unclear7119%5.9
Project too difficult5214%6.2
Missing or damaged items4412%4.1
Delivery timing3810%6.8
More variety3510%7.4
Packaging waste298%7.9
Pause or skip226%5.2

The shares add up to more than 100% because comments can carry several codes. That is fine, as long as the report says so.

Work out the recommend score yourself too. With 0 to 10 questions the usual convention is that 9 and 10 are promoters, 7 and 8 are passives and 0 to 6 are detractors; the Net Promoter Score is the percentage of promoters minus the percentage of detractors. Here 165 of 412 people scored 9 or 10 (40%) and 99 scored 0 to 6 (24%), so the score is 40 minus 24, which gives 16. A chat assistant will happily produce this number too, but it is a two-cell formula and there is no reason to trust anything else with it.

The cross-tab is where the insight usually sits. Filter the tags to detractors only and the picture changes: 29 of the 99 detractors mentioned missing or damaged items. That theme is only 12% of all comments, yet nearly three in ten unhappy customers raised it. A fix to packing checks would move the score more than any change to the craft itself.

Spot-check the tags before anyone sees a chart

Pick 40 tagged comments at random, hide the AI's codes, tag them yourself, then compare. If you disagree on more than about one in ten, the codebook needs work rather than the AI needing a better mood. In the craft-box check, the owner disagreed on 6 of 40 (15%). Five of the six were the same confusion: comments like "too fiddly for me, the diagram for step 4 was tiny" had been coded as "too difficult" when the customer was plainly blaming the diagram. She added a rule, "if a specific instruction step or diagram is named, use instructions unclear", re-ran the two affected batches, and the next check showed 3 disagreements in 40.

Watch for these patterns while you check:

  • Sarcasm read as praise. "Brilliant, another month of pom-poms" was tagged positive in the first run. The rule in the prompt fixed most of these.
  • Everything drifting into one theme. If "loved the project" swallows comments that only say "fine, thanks", add a neutral theme or leave them uncoded.
  • A bloated Other bucket. If T98 holds more than about 8% of comments, read them; there is usually a theme you missed. Here it was "wants a video tutorial", which appeared 19 times.
  • Batch drift. Compare theme shares between the first and last batch. A theme that jumps from 5% to 20% between batches usually means the model's reading changed, not your customers.

Pull real quotes for every big number

People act on quotes more than on percentages, so the report needs a few verbatim comments beside each theme. Do not ask the AI to "give some example quotes": it may tidy the grammar, blend two comments, or produce a sentence nobody wrote. Instead, filter the tags tab by theme, pick five response IDs, and copy the comment text straight from the original column. If you want help choosing, ask the assistant for "the five response IDs that best represent this theme" and then fetch the text yourself.

A quick before and after shows why this matters. The model's summary line for missing items read: "Customers frequently reported missing components, especially thread and needles, and felt let down." The comments behind it said something more specific: nine of the 44 named the same missing item, the blunt wool needle, and six mentioned they had received the box as a gift, so the missing piece embarrassed them in front of someone else. That second detail changes the fix: a spare needle taped inside the lid costs pennies and protects the gift experience.

Turn the themes into decisions with owners

A theme table is not a plan. Finish with a short decision sheet that the team can argue about, filled in like this:

FindingDecisionOwnerHow we'll know next survey
29 of 99 detractors mention missing itemsAdd a packing checklist and a spare needle to every boxWarehouse leadMissing-items theme under 5%
71 comments on unclear instructions, step diagrams named mostTest instructions on two new customers before printProduct designerInstructions theme under 10%
19 people asked for a video tutorialFilm one short video for box 14 as a trialOwnerCount mentions of the video
22 want to pause or skipMake the skip option visible in the account pageOwnerSkips versus cancellations next quarter

Rank the rows by three things: how many customers raised it, how much it hurts (the average score of people mentioning it is a good proxy), and how cheap it is to fix. The spare needle wins on all three.

Different surveys need slightly different handling

The same routine adapts to most feedback formats, but the codebook and the checks change.

  • Cancellation surveys. The subscription box's exit form asks one question, "Why are you leaving?", and the answers are short and blunt. Code the stated reason and a separate "could we have saved this?" flag. Combine this with billing data, as in running a churn analysis, because stated reasons and actual behaviour often differ.
  • Post-purchase surveys in a shop. A sports equipment shop that emails a two-question survey after each purchase gets comments tied to product types. Add the product category as a column and pivot themes by category. A theme like "sizing advice was wrong" matters far more if it clusters on running shoes than if it is spread evenly across the shop.
  • Event or programme feedback. A community interest company running weekend repair cafés collects 30 to 40 paper forms per event. At that size, typing up the forms is the slow part, and the AI is most useful for reading handwriting, as covered in turning handwritten forms into spreadsheet data. The tagging can then be done by hand in twenty minutes.
  • Complaints and returns notes. These are not surveys, but the tagging method is identical and the counts often explain survey results. The joined-up approach is in spotting patterns in complaints, returns and defects.

How long a 400-response survey takes

For the craft-box survey, the first run took the owner about three and a half hours: 20 minutes to clean the export, 40 minutes on the codebook, 30 minutes of tagging in batches, 45 minutes of spot-checking and one re-run, and about an hour and a half on the pivot tables, quotes and decision sheet. The second survey took about 90 minutes, because the codebook and pivot layout already existed.

The cost is whatever AI plan you have. A $20-a-month chat plan handles this volume comfortably. If you run large surveys every month, tagging through an automation with an API model is cheaper per comment, but for most small firms a monthly or quarterly survey does not justify building one.

Before uploading anything, check which plan you are on. Business plans such as ChatGPT Business and Claude Team do not train on your content by default. On individual plans, switch off the model-training setting in privacy settings first. Stripped of names and contact details, most survey comments are low risk, but the stray personal details mentioned earlier are not.

Keep the codebook fixed so surveys can be compared

The real value arrives with the second and third surveys, when you can see whether the instructions theme fell after you changed the diagrams. That only works if the codes mean the same thing each time. Save the codebook as a versioned document (v1.0, v1.1) with the definitions and rules, and change it deliberately:

  • New themes go into Other with a note first. Promote one to its own code only if it appears in two surveys running.
  • When you split or merge a theme, record the date and re-tag the previous survey's comments for that theme, or mark the trend line as broken at that point.
  • Keep the same questions, in the same order, sent at the same point in the customer's month. Changing the question wording changes the answers more than most owners expect.

The same discipline applies inside the business: the method in analysing staff surveys and exit interviews uses a fixed codebook too, and staff themes often explain customer ones. When the craft box's packing team said in their own survey that the bench was too small for the new box size, the missing-items problem made a lot more sense.

Survey analysis questions owners ask next

How many responses do I need before AI analysis is worth it?

Below about 50 written comments, read them yourself in one sitting; it takes less time than setting up a codebook and you will notice nuance a tag list misses. From roughly 100 comments upwards the tagging method pays off, and from 300 it is the only way most owners will finish the job before the next survey goes out.

Can AI tag comments written in several languages?

Yes. Current chat assistants tag comments in most widely used languages against an English codebook. Ask it to keep the original wording in your quotes file and translate only for the report. Have a fluent speaker spot-check ten tags in your biggest non-English group, because sarcasm and politeness conventions are where cross-language tagging slips.

Should I just use the AI summary built into my survey tool?

Built-in theme summaries are fine for a first look. Before relying on one, check that you can click a theme and see exactly which responses sit under it, and that you can export those tags to a spreadsheet. If you cannot audit or export the tags, treat the summary as a rough impression rather than a result.

Is a sentiment score per comment useful?

On its own, rarely. A positive or negative label tells you little that the rating question has not already told you. Sentiment becomes useful when you combine it with themes, for example finding that comments about packaging are mostly positive while comments about instructions are mostly negative. Always pair it with a theme code.

Further reads

Sources: Google Docs Editors Help, Use the AI function in Google Sheets; Microsoft Support, COPILOT function; OpenAI and Anthropic plan pages for ChatGPT Plus and Claude Pro pricing.

Want your survey comments tagged the same way every time?

On a 1:1 call we'll look at the surveys you already run, build a codebook for your business, and set up a tagging and counting routine your team can repeat after every survey.

Book a 1:1 call with me