Export the responses to a spreadsheet, remove names and emails, and ask AI to draft a short list of themes from a sample of about 50 comments. Then tag every comment against that fixed list one row at a time, count the tags with a pivot table, and hand-check about one comment in ten before acting on anything.
The usual mistake is pasting 400 answers into a chat and asking what customers think. You get a fluent summary with percentages the model estimated rather than counted, and quotes that are sometimes stitched together from two or three people. Keep the jobs apart: AI reads and labels, the spreadsheet counts, and you decide what the numbers mean.
A craft-kit subscription box and its 412 survey replies
The method is easiest to follow on one business, so the rest of this page uses an illustrative case. A monthly craft-kit subscription box with about 2,100 subscribers emails a short survey after its twelfth box. It gets 412 responses in a week. The survey has four questions:
- How likely are you to recommend us to a friend? (0 to 10)
- What did you think of this month's box? (free text)
- If you could change one thing, what would it be? (free text)
- Did you finish the project? (yes, partly, not yet)
The two free-text questions produce 366 usable comments; the rest are blank or say "n/a". Reading all 366 would take the owner most of a day and she would still have no counts at the end. The goal is a table showing which themes come up, how often, and how they relate to the recommend score, so she can decide what to change for box fourteen.
Shape the export so the AI can read it without guessing
Every survey tool can export to CSV or straight into Google Sheets or Excel. Before any AI sees the file, rearrange it into one row per response with these columns:
- response_id: a short code such as R0001. If your tool's IDs are long strings, make your own. Every tag and every quote will point back to this ID.
- score: the 0 to 10 answer, as a number.
- comment_box and comment_change: the two free-text answers, untouched.
- finished: the multiple-choice answer.
- plan or tenure: anything you already know about the customer that might explain a pattern, such as months subscribed.
Delete the name, email and address columns, and keep the original export somewhere else. Then skim the comments for personal details people typed into the text boxes. In this survey one parent wrote about her son's diagnosis and why the fine motor work suited him; another gave a phone number and asked for a callback. Those details belong in a follow-up email, not in an AI tool. If you have more than a few, the routine in redacting personal data before sharing documents saves time.
One more change pays off later. If people answered two free-text questions, keep them in separate columns rather than joining them. "What would you change" answers are nearly all complaints or wishes, while "what did you think" answers lean positive, and merging them makes every theme look more mixed than it is.
Draft a codebook from 50 comments before tagging anything
A codebook is simply the fixed list of themes you will tag against, each with a code and a one-line definition. Build it from a random sample, not the first 50 rows, because early responders are often your keenest customers. Sort the sheet by a random number column, copy 50 comments, and use a prompt like this:
Below are 50 comments from a customer survey for a monthly craft-kit subscription box.
Propose a codebook of 8 to 12 themes that together cover most of these comments.
For each theme give: a code (T01, T02...), a short name, a one-sentence definition,
and the response IDs from this sample that fit it.
Rules:
- Themes must be about one thing each (not "delivery and packaging").
- Keep complaints and praise about the same topic as separate themes.
- Add T98 Other and T99 Unclear.
- Do not invent themes that no comment supports.
Comments:
R0187 | The instructions for the macrame bit made no sense, gave up halfway
R0033 | Loved it, finished in one evening, my daughter wants the next one
...
An illustrative first draft back from the assistant looked like this:
T01 Project too difficult - customer found the project hard to complete (R0187, R0290...)
T02 Loved the project - positive about the craft itself (R0033, R0145...)
T03 Missing or damaged items - something absent, short or broken on arrival
T04 Delivery timing - box arrived late or on an awkward day
T05 Value - comments about price or what you get for it
T06 More variety - wants different crafts or fewer repeats
T07 Packaging - comments on amount or type of packaging
T08 Pause or skip - wants to skip a month or pause
T09 Suitable for children - comments about using the kit with kids
T98 Other
T99 Unclear
This is a decent start, and it has three problems you should expect in any first draft. T01 mixes two causes: some people found the project hard, others found the instructions unclear, and those need different fixes. T05 hides two opposite complaints, "too expensive" and "not enough yarn to finish". T09 is a topic, not a theme; half those comments were praise and half said the kit was not suitable for under-tens. After editing, the working codebook had twelve themes: instructions unclear, project too difficult, project too easy, loved the project, missing or damaged items, not enough materials, price too high, delivery timing, more variety, packaging waste, pause or skip, and child suitability, plus Other and Unclear. Write a one-line definition for each and a rule for any pair people could confuse.
Tag every comment against the fixed list
Now the AI does the slow part: reading each comment and choosing codes from your list. You have three reasonable ways to run it, depending on the tools you have.
Option 1: a chat assistant with file upload
ChatGPT Plus and Claude Pro both cost about $20 a month and both accept a spreadsheet upload. Send 80 to 100 comments per message rather than all 366 at once; accuracy drifts on long batches and you want to be able to re-run one batch without redoing the lot. Use the same prompt every time:
Tag each survey comment below using ONLY this codebook: [paste codebook with definitions].
A comment can have up to 3 codes. Rules:
- Tag what the customer says, not what you guess they feel.
- If the customer blames the instructions, use T01 (instructions unclear), not T02.
- Sarcasm is negative: "Great, glue again" = negative.
- If nothing fits, use T98 Other and add a 3-word note.
Return CSV only, one line per comment, in the same order:
response_id,codes,sentiment
Sentiment must be one of: positive, negative, mixed, neutral.
Illustrative output for five rows:
R0187,T01,negative
R0033,T04;T12,positive
R0211,T05;T06,negative
R0302,T98 (wants video tutorial),neutral
R0350,T10,positive
Paste each batch's CSV into a new "tags" tab. Check that the row count matches the batch before you move on; a missing row usually means the model merged two short comments.
Option 2: the AI function inside Google Sheets
If your Google Workspace or Google AI plan includes it, the AI function in Google Sheets lets you put the tagging prompt in a formula next to each comment, such as =AI("Using this codebook: ... return only the codes for this comment", C2). Google's help page says only the first 350 selected cells with AI functions generate at a time and that generation limits apply, so run the column in chunks and paste the results as values once they settle, or they may regenerate differently later. For more prompt patterns inside a sheet, see these Google Sheets AI prompts.
Option 3: Copilot in Excel
Microsoft withdrew the COPILOT worksheet function on 14 September 2026, and cells that used it now show a #NAME? error when they recalculate, according to Microsoft's COPILOT function page. The Copilot pane in Excel can still classify a text column and add the results as a new column. Give it the same codebook and rules, and check a sample exactly as you would with the other options.
Count in the spreadsheet, never in the chat
Once every comment has codes, split the codes into one row per code (a comment with two codes becomes two rows) and build a pivot table: themes down the side, count of response IDs, and the average recommend score of the people who mentioned each theme. The illustrative result for the craft box:
| Theme | Comments | Share of 366 commenters | Average recommend score |
|---|---|---|---|
| Loved the project | 118 | 32% | 8.9 |
| Instructions unclear | 71 | 19% | 5.9 |
| Project too difficult | 52 | 14% | 6.2 |
| Missing or damaged items | 44 | 12% | 4.1 |
| Delivery timing | 38 | 10% | 6.8 |
| More variety | 35 | 10% | 7.4 |
| Packaging waste | 29 | 8% | 7.9 |
| Pause or skip | 22 | 6% | 5.2 |
The shares add up to more than 100% because comments can carry several codes. That is fine, as long as the report says so.
Work out the recommend score yourself too. With 0 to 10 questions the usual convention is that 9 and 10 are promoters, 7 and 8 are passives and 0 to 6 are detractors; the Net Promoter Score is the percentage of promoters minus the percentage of detractors. Here 165 of 412 people scored 9 or 10 (40%) and 99 scored 0 to 6 (24%), so the score is 40 minus 24, which gives 16. A chat assistant will happily produce this number too, but it is a two-cell formula and there is no reason to trust anything else with it.
The cross-tab is where the insight usually sits. Filter the tags to detractors only and the picture changes: 29 of the 99 detractors mentioned missing or damaged items. That theme is only 12% of all comments, yet nearly three in ten unhappy customers raised it. A fix to packing checks would move the score more than any change to the craft itself.
Spot-check the tags before anyone sees a chart
Pick 40 tagged comments at random, hide the AI's codes, tag them yourself, then compare. If you disagree on more than about one in ten, the codebook needs work rather than the AI needing a better mood. In the craft-box check, the owner disagreed on 6 of 40 (15%). Five of the six were the same confusion: comments like "too fiddly for me, the diagram for step 4 was tiny" had been coded as "too difficult" when the customer was plainly blaming the diagram. She added a rule, "if a specific instruction step or diagram is named, use instructions unclear", re-ran the two affected batches, and the next check showed 3 disagreements in 40.
Watch for these patterns while you check:
- Sarcasm read as praise. "Brilliant, another month of pom-poms" was tagged positive in the first run. The rule in the prompt fixed most of these.
- Everything drifting into one theme. If "loved the project" swallows comments that only say "fine, thanks", add a neutral theme or leave them uncoded.
- A bloated Other bucket. If T98 holds more than about 8% of comments, read them; there is usually a theme you missed. Here it was "wants a video tutorial", which appeared 19 times.
- Batch drift. Compare theme shares between the first and last batch. A theme that jumps from 5% to 20% between batches usually means the model's reading changed, not your customers.
Pull real quotes for every big number
People act on quotes more than on percentages, so the report needs a few verbatim comments beside each theme. Do not ask the AI to "give some example quotes": it may tidy the grammar, blend two comments, or produce a sentence nobody wrote. Instead, filter the tags tab by theme, pick five response IDs, and copy the comment text straight from the original column. If you want help choosing, ask the assistant for "the five response IDs that best represent this theme" and then fetch the text yourself.
A quick before and after shows why this matters. The model's summary line for missing items read: "Customers frequently reported missing components, especially thread and needles, and felt let down." The comments behind it said something more specific: nine of the 44 named the same missing item, the blunt wool needle, and six mentioned they had received the box as a gift, so the missing piece embarrassed them in front of someone else. That second detail changes the fix: a spare needle taped inside the lid costs pennies and protects the gift experience.
Turn the themes into decisions with owners
A theme table is not a plan. Finish with a short decision sheet that the team can argue about, filled in like this:
| Finding | Decision | Owner | How we'll know next survey |
|---|---|---|---|
| 29 of 99 detractors mention missing items | Add a packing checklist and a spare needle to every box | Warehouse lead | Missing-items theme under 5% |
| 71 comments on unclear instructions, step diagrams named most | Test instructions on two new customers before print | Product designer | Instructions theme under 10% |
| 19 people asked for a video tutorial | Film one short video for box 14 as a trial | Owner | Count mentions of the video |
| 22 want to pause or skip | Make the skip option visible in the account page | Owner | Skips versus cancellations next quarter |
Rank the rows by three things: how many customers raised it, how much it hurts (the average score of people mentioning it is a good proxy), and how cheap it is to fix. The spare needle wins on all three.
Different surveys need slightly different handling
The same routine adapts to most feedback formats, but the codebook and the checks change.
- Cancellation surveys. The subscription box's exit form asks one question, "Why are you leaving?", and the answers are short and blunt. Code the stated reason and a separate "could we have saved this?" flag. Combine this with billing data, as in running a churn analysis, because stated reasons and actual behaviour often differ.
- Post-purchase surveys in a shop. A sports equipment shop that emails a two-question survey after each purchase gets comments tied to product types. Add the product category as a column and pivot themes by category. A theme like "sizing advice was wrong" matters far more if it clusters on running shoes than if it is spread evenly across the shop.
- Event or programme feedback. A community interest company running weekend repair cafés collects 30 to 40 paper forms per event. At that size, typing up the forms is the slow part, and the AI is most useful for reading handwriting, as covered in turning handwritten forms into spreadsheet data. The tagging can then be done by hand in twenty minutes.
- Complaints and returns notes. These are not surveys, but the tagging method is identical and the counts often explain survey results. The joined-up approach is in spotting patterns in complaints, returns and defects.
How long a 400-response survey takes
For the craft-box survey, the first run took the owner about three and a half hours: 20 minutes to clean the export, 40 minutes on the codebook, 30 minutes of tagging in batches, 45 minutes of spot-checking and one re-run, and about an hour and a half on the pivot tables, quotes and decision sheet. The second survey took about 90 minutes, because the codebook and pivot layout already existed.
The cost is whatever AI plan you have. A $20-a-month chat plan handles this volume comfortably. If you run large surveys every month, tagging through an automation with an API model is cheaper per comment, but for most small firms a monthly or quarterly survey does not justify building one.
Before uploading anything, check which plan you are on. Business plans such as ChatGPT Business and Claude Team do not train on your content by default. On individual plans, switch off the model-training setting in privacy settings first. Stripped of names and contact details, most survey comments are low risk, but the stray personal details mentioned earlier are not.
Keep the codebook fixed so surveys can be compared
The real value arrives with the second and third surveys, when you can see whether the instructions theme fell after you changed the diagrams. That only works if the codes mean the same thing each time. Save the codebook as a versioned document (v1.0, v1.1) with the definitions and rules, and change it deliberately:
- New themes go into Other with a note first. Promote one to its own code only if it appears in two surveys running.
- When you split or merge a theme, record the date and re-tag the previous survey's comments for that theme, or mark the trend line as broken at that point.
- Keep the same questions, in the same order, sent at the same point in the customer's month. Changing the question wording changes the answers more than most owners expect.
The same discipline applies inside the business: the method in analysing staff surveys and exit interviews uses a fixed codebook too, and staff themes often explain customer ones. When the craft box's packing team said in their own survey that the bench was too small for the new box size, the missing-items problem made a lot more sense.
Survey analysis questions owners ask next
How many responses do I need before AI analysis is worth it?
Below about 50 written comments, read them yourself in one sitting; it takes less time than setting up a codebook and you will notice nuance a tag list misses. From roughly 100 comments upwards the tagging method pays off, and from 300 it is the only way most owners will finish the job before the next survey goes out.
Can AI tag comments written in several languages?
Yes. Current chat assistants tag comments in most widely used languages against an English codebook. Ask it to keep the original wording in your quotes file and translate only for the report. Have a fluent speaker spot-check ten tags in your biggest non-English group, because sarcasm and politeness conventions are where cross-language tagging slips.
Should I just use the AI summary built into my survey tool?
Built-in theme summaries are fine for a first look. Before relying on one, check that you can click a theme and see exactly which responses sit under it, and that you can export those tags to a spreadsheet. If you cannot audit or export the tags, treat the summary as a rough impression rather than a result.
Is a sentiment score per comment useful?
On its own, rarely. A positive or negative label tells you little that the rating question has not already told you. Sentiment becomes useful when you combine it with themes, for example finding that comments about packaging are mostly positive while comments about instructions are mostly negative. Always pair it with a theme code.
Further reads
- Can AI Analyse My Sales Spreadsheet? What to Upload and Check — What to upload and check when AI works on your sales figures.
- What Are Diners Really Saying? Using AI to Analyse Your Reviews — The same theme-counting idea applied to public reviews.
- How to Turn Customer Reviews Into Marketing Copy With AI — Turn the praise your survey uncovered into honest marketing copy.
- How to Build a KPI Dashboard With AI When You Have No Data Team — Track your survey scores alongside the rest of your numbers.
- How to Set Up Human Review for AI Work Without Slowing Down — A light review routine for any AI output you act on.
- How to Summarise Long Documents With AI Without Missing Details — Summarise long reports without losing the details that matter.
- How to Measure Customer Reaction After Introducing AI — Four signals, survey wording that doesn't lead, a conversation-sorting prompt and a decision rule for reading small-business numbers honestly.
- How to Map Your Customer Journey and Find Where AI Helps — A journey grid, evidence sources and the signs a step suits AI, walked through with a new puppy owner's first year at a veterinary practice.
- How to Win Back Lapsed Customers With AI-Personalised Emails — Plan a restrained win-back campaign using verified customer history, two useful emails and a comparison with customers you leave uncontacted.
- How to Test a New Service Idea With AI Before You Launch It — A six-week way to test a new service with AI doing the drafting and real customers supplying the evidence, with pass marks set before the results arrive.
- How to Use AI to Find Out Why Website Visitors Don't Buy — Combine your analytics funnel, free session recordings and exit-survey answers, then have AI rank why visitors leave, with an online bookshop worked through.
- What Is Review Gating and Can It Get Your Business Penalised? — Check whether your feedback forms, AI sentiment labels or review reminders quietly exclude unhappy customers.
- AI Social Listening: Tracking What Customers Say About You — Build a small listening process that separates real customer themes from duplicate posts, wrong matches and misleading scores.
- Are AI-Generated Customer Personas Accurate? How to Check — Why AI personas read convincingly and still miss, plus a two-hour audit with source tagging, data checks, recognition calls and prompts that show their working.
- Hire a Data Analyst or Use AI for Your Reporting? — A bookshop's reporting problem priced out: when AI tools run by your own team are enough, when a freelance analyst should build the base, and when to hire.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Google Docs Editors Help, Use the AI function in Google Sheets; Microsoft Support, COPILOT function; OpenAI and Anthropic plan pages for ChatGPT Plus and Claude Pro pricing.