Put your outputs, outcome measures and feedback into one tidy spreadsheet, ask an AI assistant to calculate the figures and show its working, check each number against the source, then have it draft the report section by section around a fixed outline: the need, what you did, what changed, what you learned. Consented stories go in last.
The trap is in the middle step. AI is good at arithmetic done in code, at sorting hundreds of comments into themes, and at turning verified figures into readable paragraphs. It has no idea what your numbers mean. Give it a survey where 40% of respondents ticked "less lonely" and it will cheerfully write "our programme reduced loneliness by 40%". Deciding what you delivered, what changed, and what you can honestly say caused it is the charity's job, and the part funders read most closely.
Outputs, outcomes and what you can honestly claim
Three words get muddled in almost every first draft, so settle them before you open a chat window. Throughout this tutorial the illustration is a small charity running physiotherapist-led strength and balance classes for older people at risk of falls.
| Term | Meaning | Example from the falls charity | Safe wording |
|---|---|---|---|
| Output | What you delivered | 14 groups, 336 classes, 186 people enrolled | "We ran 336 classes" |
| Outcome | What changed for people | Balance test scores at week 1 and week 12 | "Of the 121 people tested twice, 84 improved" |
| Impact | The longer-term difference you contributed to | Fewer falls, staying independent at home | "Participants reported fewer falls; we can't measure how many would have fallen anyway" |
Most small charities can report outputs with certainty, outcomes with care, and impact only as what people told them. A report that is honest about that difference reads as more credible, not less.
Getting everything into one sheet an AI can read
AI tools struggle with the way charity data is usually kept: a register per group, a survey in a separate form tool, case notes in documents. Spend your time here and the rest goes quickly. Build one sheet, one row per participant, with columns like these:
- participant_id (never a name), group, start_date, age_band (for example 70-79, rather than date of birth)
- classes_offered and classes_attended
- balance_wk1 and balance_wk12, left blank if not tested, never zero
- confidence_wk1 and confidence_wk12 from your short questionnaire
- falls_reported_before and falls_reported_during, as participants reported them
- comment: the free-text answer to "What difference, if any, have the classes made?"
Add a second tab called "data dictionary" with one line per column saying what it means and how it was collected. The AI reads it, and so will next year's staff. A few lines from the falls charity's version:
balance_wk1 Balance score at first class, 0-56, HIGHER is better.
Taken by the physiotherapist. Blank = not tested.
walk_time_wk1 Seconds to walk 4 metres at normal pace. LOWER is better.
confidence_wk1 Sum of 5 questions scored 1-4 (range 5-20), HIGHER is
better. Self-completed; volunteers helped 9 people.
classes_offered Classes scheduled for that person's group (24 unless
the group started late).
The direction words matter more than they look. In a first run without them, the assistant treated a rise in walking time as an improvement, because for every other measure a higher number was better, and reported that 70% of participants had "improved" their walking speed. A spreadsheet filter showed most had actually got faster, which meant lower times; the figure was close to the reverse of the truth. Adding "LOWER is better" to the dictionary and asking again gave the correct count. Then, before any calculating, ask the assistant to audit the data:
This sheet has one row per participant in our falls-prevention classes.
The "data dictionary" tab explains each column. Before calculating anything,
check the data and list:
- duplicate participant_ids, or the same person appearing to enrol twice
- rows where week-12 scores exist but week-1 scores don't (and the reverse)
- attended greater than offered, or impossible values
- any column where blanks and zeros are mixed
Don't fix anything. Give me the row numbers so I can check them.
A typical reply (illustrative): "Rows 44 and 139 share participant_id P0521 with different start dates. 17 participants have a week-12 balance score but no week-1 score. Row 88 shows 26 classes attended out of 24 offered. The confidence columns use 0 for 11 people whose other scores are blank, which may mean 'not asked' rather than a score of zero." Each of these would have skewed the results. The last one is the classic: a zero that means "missing" drags an average down and makes the programme look worse than it was.
Calculating with the working on show
Once the sheet is clean, ask for figures in a way that forces the assistant to show how many people each one is based on. This is the single habit that makes an AI-assisted report trustworthy.
Calculate the following. For every figure, state n (how many participants it's
based on) and show the code you used.
1. Enrolled, completed (attended 18 or more of 24 classes), average attendance
2. For participants with BOTH week-1 and week-12 balance scores:
how many improved, stayed the same, got worse; median change
3. The same for confidence scores
4. Falls reported before vs during, for people who answered both questions
Use only complete pairs for before/after figures. Don't round until the end.
Put all figures in one table I can paste into a "verified figures" sheet.
Illustrative output:
| Measure | Result | n |
|---|---|---|
| Enrolled | 186 | 186 |
| Completed (18+ of 24 classes) | 131 (70%) | 186 |
| Balance improved, week 1 to 12 | 84 of 121 (69%) | 121 |
| Confidence improved | 96 of 118 (81%) | 118 |
| Reported one or more falls during the programme | 19 of 124 (15%) | 124 |
Now recheck two or three figures yourself with a spreadsheet filter. If the assistant says 84 improved, filter the sheet for week-12 score higher than week-1 score and count. When a figure doesn't match, ask the assistant why; usually it has read a column differently from the way your dictionary defines it. Only figures that have passed this check go into the verified figures sheet, and the report draws on nothing else. The reason for all this caution is covered in why AI makes things up and how to catch it: a model that runs out of data fills the gap fluently. If you're working inside Excel or Sheets, Copilot in Excel and Gemini in Sheets can do the same calculations in the file; the n-and-check habit is the same. Being able to spot made-up figures in AI-drafted documents pays off here too.
Finding themes in comments without cherry-picking
Free-text comments are where impact reports come alive, and where they're most often distorted: the three glowing quotes go in, the complaints about the church hall's heating don't. Ask the AI to theme all of them, including the uncomfortable ones.
Below are 142 participant comments, one per line, each starting with a participant_id.
Group them into 5-8 themes. For each theme give: a name, the number of comments,
two quotes that represent it fairly, and one quote that is an exception.
Include negative and neutral themes. Don't combine or reword quotes.
A reply might list "confidence going out alone" (48 comments), "social contact, meeting people" (39), "specific physical gains such as getting up from a chair" (31), "wanted more classes or a longer programme" (17) and "venue or transport problems" (11). Read at least 20 comments yourself to check the themes hold, then report them with counts: "48 of 142 people mentioned feeling more confident going out alone" is far stronger than a single hand-picked quote. The "wanted more classes" theme belongs in the report too. Funders like to see that you listen.
Add those theme counts up and they come to 146, four more than the 142 comments. That's usually fine, because a comment such as "I've made friends and I can get out of a chair on my own now" belongs in two themes, but it needs saying in the report ("some comments appear under more than one theme") or a careful reader will think the arithmetic is wrong. It's also the moment to ask the assistant for the participant_ids behind each theme. With the IDs in hand you can open ten comments from the largest theme and check they really are about confidence going out alone, rather than confidence in general, which is the kind of quiet widening that inflates a headline count.
Drafting the report to a fixed outline
Give the AI the outline, the verified figures sheet, the theme summary and any consented case study, and have it draft one section at a time. Drafting the whole report in one go invites filler and invented connecting claims.
| Section | What goes in | Source | Length |
|---|---|---|---|
| The need | Why falls prevention matters for the people you serve | Your referral data; one cited public source | 150 words |
| What we did | Groups, classes, who came | Verified figures (outputs) | 200 words |
| What changed | Before-and-after results with n | Verified figures (outcomes) | 250 words |
| What people said | Themes with counts, fair quotes | Theme summary | 200 words |
| One person's story | A consented case study | Interview notes | 200 words |
| What we learned | What worked, what didn't, what's changing | Staff discussion | 150 words |
Include this instruction with every section: "Use only the figures and quotes provided. Where a sentence claims the programme caused something, rewrite it as what participants showed or reported." Then read for overclaiming anyway. A before and after from the "what changed" section:
Before: Thanks to our classes, 69% of participants significantly improved their balance, dramatically reducing their risk of falling.
After: Of the 121 people who took the balance test at both the start and end, 84 (69%) scored better at week 12. Fewer participants reported falls during the programme than in the three months before it, though we can't say how many falls the classes prevented.
"Significantly" implies a statistical test nobody ran, and "dramatically reducing their risk" is a clinical claim the data can't support. The after version is longer and reads as more convincing to anyone who knows the field.
If you keep past reports, session plans and interview notes as documents, Gemini Notebook (formerly NotebookLM) is useful for the "what we learned" section: it answers from the sources you give it and shows which document each point came from.
When your data looks nothing like the falls charity's
Few charities have a tidy week-1 and week-12 test. The method still works; what changes is which rows of the outline you can fill with numbers.
- A youth mentoring charity with 14 volunteer mentors and 60 young people might have no test at all, just mentor session notes and an end-of-year conversation. Paste anonymised session notes (with the young person's ID, not name) and ask the AI to count how often goals were set, reviewed and met, with quotes. Report those as "goals met, as recorded by mentors", and say that's what they are.
- A food bank is mostly outputs: parcels, households, repeat visits, referrals by source. The useful AI job is spotting patterns across months ("repeat visits rose from 22% to 31% of households between spring and autumn") and drafting the need section from your referral data. Resist letting the AI turn parcels into "meals provided" unless you've defined how many meals a parcel contains.
- A charity-run nursery reporting to a funder can combine attendance, places funded for families on low incomes, and practitioners' development notes. Here the identification risk is higher still: children, small groups and a known local setting. Aggregate to the whole setting and never quote a child's words with details attached.
If a figure can't be supported, leave it out and say what you'll measure next year. That sentence costs nothing and funders rarely hold it against you.
Small numbers and people who could be recognised
Impact reports are public, and small charities work with small groups. A breakdown such as "the two participants over 90 in the Thursday group both improved" identifies people. My rule: don't publish any figure based on fewer than five people, and don't combine details (age band, group, condition, area) that together narrow down to an individual. Ask the AI to flag any cell in your figures table with n under 5 before you draft. For the falls charity, the flag came back on three cells: the 90-and-over age band in two groups (n of 2 and 3) and one group's falls figure (n of 4). The fix was to merge age bands into "80 and over" and report falls for all groups together, which cost nothing in meaning. A case-study draft needed the same care. The assistant turned interview notes into "a former headteacher who lives alone on the estate behind the community centre", which anyone in that neighbourhood could identify. The published version said "a woman in her eighties who lives alone", with the rest of her story unchanged and her consent form covering it.Case studies need written consent covering the report and any other channel where you'll reuse the story, and names changed unless the person has asked to be named.
A funder's-eye check before you send it
Read the finished draft once as the funder would, with this list beside you:
- Every figure appears in the verified figures sheet, with its n somewhere nearby.
- Percentages are never shown without the number of people behind them.
- Outputs, outcomes and reported impact are labelled as such.
- Claims about cause are worded as what people showed or said.
- Last year's figures, if compared, were calculated the same way.
- At least one limitation or thing that didn't work is included.
- Quotes are exact, and every story has consent on file.
Run against the falls charity's first full draft, the list found two failures. Item 2: the "what people said" section read "81% felt more confident", with no n; it became "96 of the 118 people who answered both questionnaires (81%) scored higher for confidence at week 12". Item 5 was more serious. The draft compared this year's 70% completion with "74% last year", but last year's report had counted anyone attending 16 of 24 classes as completing, and this year's threshold was 18. Recalculated at 18, last year's figure was 66%. The honest line became "Completion rose from 66% to 70% on the same definition", which is a smaller claim, but the right direction. Had the funder spotted the mismatch themselves, every other figure in the report would have been doubted.
A useful final prompt, once the draft is finished, is to ask the assistant to act as a sceptical reviewer: "List every sentence in this report that makes a claim not supported by the verified figures sheet or the theme summary, quoting each one." It won't catch everything, and it sometimes flags fair sentences, but it's a quick second pass before a person does the final read.
For the falls charity, the whole process might take about two days the first time: a day building and cleaning the sheet, half a day calculating and checking, half a day drafting and editing. The second year is quicker, because the sheet, dictionary and outline already exist. If you want to reuse the same figures in a bid, writing grant applications with AI picks up from here; just keep the verified figures sheet as the single source for both.
Impact report questions
Can AI make the charts for an impact report?
Yes. ChatGPT and Claude can both produce charts from an uploaded spreadsheet, and Copilot in Excel or Gemini in Sheets can build them inside the file. Ask for simple bar or line charts with the number of people behind each figure labelled. Check each chart against your verified figures sheet, because a chart drawn from the wrong column looks just as convincing as a right one.
Should we tell funders that AI helped write the report?
If the funder asks, answer honestly: AI helped calculate and draft, and staff checked every figure and wrote the final text. Some funders have their own guidance on AI in applications and reports, so read their terms. What funders care about most is that the numbers are true and traceable, which the checking steps in this tutorial make possible.
What if we never collected before-and-after measures?
Report what you do have, honestly labelled: outputs such as sessions and attendance, plus what participants said, themed fairly. Don't let AI turn comments into an outcome percentage. Then fix the gap for next year by adding one short measure at the start and end of your programme. Even a three-question survey gives you a real before and after to report.
Further reads
- AI Grant Writing Tools Compared for Small Non-Profits — Tools that help turn report findings into funding bids.
- Can You Use AI for Grant Writing Without Funders Rejecting It? — How funders view AI-written applications and reports.
- How to Automate Monthly Management Reports With AI — Build the monthly habit that makes annual reports easy.
- Hire a Data Analyst or Use AI for Your Reporting? — When the numbers need a person, not just a prompt.
- Charity Appeals and Social Posts With AI That Keep Your Voice — Turn the report's findings into appeals and posts.
- Can AI Help a Charity Forecast Demand for Its Services? — Use the same data to plan next year's demand.
- Do Small Charities Need Outside Help to Adopt AI? — What a small charity can do with AI on its own, the five signs it needs outside help, and how to brief that help so the work stays yours.
- Donor Thank-You Letters With AI That Still Feel Personal — Personal means specific, not a first-name merge: feed AI each gift's details and true impact stories, tier the response, and keep a real signature.
- How Small Charities Can Use AI With Almost No Budget — A seven-stage plan for a charity with no AI budget: get verified, pick a free workspace, choose three jobs, set a data rule, and trial it for a month.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: OpenAI and Anthropic help pages on file uploads and data analysis; Google help page on Gemini Notebook (formerly NotebookLM); Microsoft Copilot in Excel documentation.