How Tutors Should Check AI-Made Worksheets Before Using Them

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How Tutors Should Check AI-Made Worksheets Before Using Them.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How Tutors Should Check AI-Made Worksheets Before Using Them.

Check an AI-made worksheet in four passes before a pupil sees it: solve every question yourself without looking at the answer key, then compare your answers with the AI's key, then check the level against the pupil's syllabus and reading age, and finally scan diagrams, names and wording. Allow about ten minutes per sheet.

The answer key is where most errors hide. AI assistants are good at writing questions that look right and weaker at working them, especially multi-step maths, unit conversions and multiple-choice questions where two options are defensible. A neatly formatted sheet with a confident key is exactly the thing that gets waved through, which is why solving it cold comes first.

Follow me on Instagram@sagnikteaches

Why a quick glance misses AI worksheet errors

When you read someone else's worksheet, your brain checks whether each question sounds like a question. It does not work the question. AI output is fluent by design, so every item passes the "sounds right" test, and the wrong ones only reveal themselves when you pick up a pencil.

Connect on LinkedInSagnik Bhattacharya

The errors also cluster in places tutors tend to trust. The first two or three questions are usually fine. Mistakes creep in around question six onwards, where the model is combining steps, and in the answer key, which it writes after the questions and does not re-derive. A pupil who gets question eight "wrong" against a faulty key loses confidence in a method they had actually learnt, and a parent who spots it loses confidence in you.

Subscribe on YouTube@codingliquids

So the check below is built around doing the work, not reading it. It has four passes, and each pass has a small number of items with a reason and a way to verify it.

The four-pass check, item by item

Pass 1: solve it cold, with the key hidden

Cover or delete the answer section, then work every question on paper as if you were the pupil. This takes the longest and catches the most.

  • Every question can be answered from what is on the page. AI sometimes writes a comprehension question about a detail it intended to include in the passage but never did. Verify by pointing to the line in the passage that answers each question.
  • Multiple-choice questions have exactly one correct option. Look for distractors that are also true. Verify by asking of each wrong option: could a well-informed pupil argue for this?
  • The numbers behave. A sheet for an 11-year-old should not produce 7.3333 recurring unless rounding is the point. Verify that each answer is a number the pupil can reach with the methods they know.
  • Units and conventions are consistent. Mixed minutes and hours, or grams and kilograms in one question, are fine only when conversion is the skill being tested. Verify that the question tells the pupil which unit to answer in.

Pass 2: audit the answer key against your own answers

  • Every key answer matches yours. Where they differ, work it a third time before deciding who is wrong. Verify by marking disagreements in the margin, not by fixing silently.
  • Worked solutions use a method you would accept. A key that jumps straight to the answer, or uses a method the pupil has not been taught, is not much use for self-marking. Verify that each step would earn method marks in the pupil's exam.
  • Mark allocations are realistic. If the sheet imitates exam style, a two-step question worth one mark teaches the wrong habit. Verify against a real past paper's mark scheme for that topic.

Pass 3: level, syllabus and reading load

  • The content matches the pupil's syllabus. Chat assistants blend syllabuses from different exam systems. Verify by checking the topic list for the pupil's actual course, including the terms it uses (some courses say "mean", others "average").
  • The reading age fits the pupil. A maths word problem with dense wording tests reading, not maths. In Word for Microsoft 365, Home, then Editor, then Document stats shows a Flesch-Kincaid grade level, which is roughly the years of schooling a reader needs. Verify that it sits at or below the pupil's year of schooling.
  • Difficulty rises steadily and the time is right. Verify that the first questions build confidence, the last stretch the pupil, and the whole sheet fits the slot you have planned for it.

Pass 4: images, names, fairness and ownership

  • Diagrams are correct and labels are legible. AI image tools still produce garbled text, extra arrows and anatomically wrong drawings. Verify every label letter by letter, or remove the image and use a diagram you trust.
  • Names and contexts are varied and neutral. Watch for all the engineers being men, or every word problem assuming a family with a car and a garden. Verify by listing who does what across the sheet.
  • Nothing is copied. Ask yourself whether a question looks suspiciously like a well-known exam question. Verify by searching a distinctive phrase in quotation marks.
  • Spelling matches what the pupil is taught. Models default to one spelling convention. Verify "colour", "metre" or "analyse" style words match the pupil's school and exam.

Six errors that get through when nobody works the sheet

These are illustrative, but each is the kind of mistake that turns up regularly in AI-drafted practice material.

1. Successive percentages collapsed into one. Question: "A jacket costs $80. It is reduced by 15%, and then a further 10% is taken off the sale price. What is the final price?" The correct working is 80 × 0.85 = $68, then 68 × 0.9 = $61.20. An AI key that adds the discounts to 25% gives $60, which is wrong by $1.20 and teaches the exact misconception the question exists to test.

2. The "back to where you started" trap. Question: "A price rises by 20% and then falls by 20%. Is it back to the original price?" A key that says "yes" is wrong: $100 becomes $120, and 20% off $120 is $96. This one is worth keeping in your bank precisely because AI keys get it wrong.

3. Two correct options. "Which of these is a renewable energy source? A) wind B) coal C) biomass D) natural gas." Wind and biomass are both renewable. The AI key says A. A pupil who picks C has answered correctly and been marked wrong. Fix it by replacing biomass with oil.

4. An equation labelled balanced when it is not. A chemistry sheet shows "Mg + O2 → MgO" and asks pupils to identify the reactants of this balanced equation. It is not balanced: the correct form is 2Mg + O2 → 2MgO. The question is still usable, but only if you fix the equation or change the task to "balance this equation".

5. Speed with mixed units. "A cyclist travels 12 km in 40 minutes. What is her average speed in km/h?" The answer is 12 ÷ (40 ÷ 60) = 18 km/h. A key that divides 12 by 40 and reports 0.3 km/h has ignored the conversion. Worse, a key that reports 0.3 km per minute is right but in the wrong unit, and a pupil copying it learns nothing about which unit the question asked for.

6. A confident wrong date. A history or general-knowledge sheet states "the first crewed Moon landing took place in 1968". It was July 1969. Dates, names and "firsts" are exactly the details AI states fluently and gets wrong, and pupils memorise them.

A filled-in check for one worksheet

Here is how the four passes looked on an illustrative ten-question percentages sheet made for a 13-year-old preparing for end-of-year exams. The tutor asked a chat assistant for "ten mixed percentage questions, increasing difficulty, with a worked answer key".

ItemWhat the tutor foundFix
Q1-Q5 solved coldAll correct, sensible numbersNone
Q6 successive discountKey gave $60 instead of $61.20Corrected key, kept question
Q7 rise then fallKey said "back to original"Rewrote key with $100, $120, $96 working
Q8 percentage of an amountCorrect, but answer 33.3333Added "give your answer to 1 decimal place"
Q9 reverse percentageCorrect, but the key skipped the methodAdded a two-line worked solution
Q10 compound interestNot on this pupil's syllabus yetSwapped for a percentage change question
Reading loadGrade level 9.8 on Document statsShortened word problems to one sentence each
Names and contextsEvery question about shoppingChanged two to savings and a school trip budget

Result: two wrong answers in the key, one off-syllabus question and a reading load well above the pupil's level. The whole check took twelve minutes, and the sheet went from "would have caused a wobbly lesson" to reusable. The tutor saved the corrected version in a folder named by topic with "checked" and the date in the file name, so it never has to be checked again.

Using a second AI as a checker, not a judge

A second model can speed up pass two, as long as you treat its verdict as a list of things to look at rather than a ruling. Paste the questions without the key, ask for full working, then compare.

You are checking a practice worksheet for a 13-year-old.
Work each question below step by step, showing every step.
Do not assume any answer key is correct.
After your working, list any question that:
- has more than one defensible answer,
- cannot be answered from the information given,
- uses a unit the question does not ask for,
- needs a method beyond basic percentages, fractions and ratio.

Questions:
[paste questions only, no answers]

An illustrative reply on the sheet above might read: "Q6: 80 × 0.85 = 68; 68 × 0.9 = 61.2, final price $61.20. Q7: 100 × 1.2 = 120; 120 × 0.8 = 96, so the price is not back to the original. Q10 requires compound interest, which may be beyond basic percentages." That matches what the tutor found by hand. But the same reply might also claim Q4 is ambiguous when it is not, or work Q9 with a slip of its own.

The rule is simple: where your answer, the original key and the checker all agree, move on. Where any two disagree, work it by hand a third time. Never let two AI outputs outvote you. If you want a wider habit for this kind of checking, the routine in what AI hallucinations are and how to catch them applies to worksheets as well as business documents.

Ten minutes per sheet: the routine in order

  1. Minutes 0-4: solve cold. Hide the key. Work every question on paper. Mark any question you hesitate over.
  2. Minutes 4-6: compare with the key. Circle disagreements. Rework those, then correct the key.
  3. Minutes 6-8: level and reading load. Check the topic list for the pupil's course. Run Document stats or read one word problem aloud as the pupil would.
  4. Minutes 8-10: visuals, names, spelling. Read every label in every image. Skim the contexts and names. Fix spelling to the pupil's convention.
  5. Save and label. File it by topic with "checked" and the date. Note in a simple error log what the AI got wrong, because the same mistakes recur and your prompt should start preventing them.

For a sheet longer than fifteen questions, split the check across two sittings. Tired checking is where the errors get through.

Prompts that produce sheets you can check faster

You cannot prompt errors away, but you can ask for output that makes errors easier to find. The main tricks are to have the model state what each question tests, show full working, keep numbers clean and leave images out.

Create a practice worksheet.
Pupil: aged 11, second term of the year, working at expected level.
Topic: adding and subtracting fractions with different denominators.
Syllabus terms to use: "denominator", "equivalent fraction", "simplest form".
Format:
- 8 questions, increasing difficulty.
- For each question, one line in brackets saying which skill it tests.
- Denominators no larger than 12. Answers must simplify cleanly.
- No images or diagrams.
- Word problems: one sentence each, everyday contexts, varied names.
Then, on a separate section headed ANSWER KEY, give full working for every
question, step by step, ending with the answer in simplest form.
Use the spelling conventions: colour, metre, organise.

Compare the output from a loose prompt ("make a fractions worksheet for an 11-year-old") with this one. The loose version typically mixes skills, throws in mixed numbers the pupil has not met, and gives a key of bare answers. The tight version gives you a skill label to check against the syllabus and full working to check line by line, which cuts pass two roughly in half. Keeping prompts like this in a shared folder is worth it; the approach in planning a week of lessons with AI pairs well with it.

Diagrams and pictures: the part to be strictest about

Text errors can be corrected in seconds. Image errors usually cannot, because you cannot edit a generated picture reliably. A labelled plant cell with "chloroplast" spelled "chlorplast", a heart diagram with the chambers on the wrong side, or a clock face with thirteen numbers are all things AI image tools produce. A pupil who revises from a wrong diagram carries the error into the exam.

Three rules keep this manageable. First, ask for worksheets without images, then add diagrams you already trust from your own resources. Second, if you do use a generated image, zoom in and read every label, count every element, and check the orientation. Third, never use a generated image for anything a pupil must reproduce or label from memory. The broader checks in checking AI-generated images for errors cover what else to look for.

Names, contexts and fairness

AI word problems tend to drift towards the same handful of settings and the same assumptions. An illustrative sheet of ten problems might feature eight shopping trips, a family with two cars and a holiday abroad, and a "Dr Smith" who is always male. None of that is wrong, but a pupil who never sees themselves in the material notices, and some contexts quietly exclude pupils whose families do not share them.

A quick way to check: list the characters and settings in the margin. If more than half share a setting, swap two. If jobs split along obvious lines, swap the names. Prefer contexts pupils of any background can picture: a school trip budget, a bus timetable, a recipe, a sports league table. For a deeper look at the patterns to watch, see spotting bias and stereotypes in AI-written content.

When to bin the sheet and start again

Fixing a sheet is worth it when the problems are local. It is not worth it when the problems are structural. Use these thresholds:

  • More than two errors in the answer key: regenerate with a tighter prompt, or write the sheet yourself. A key with three errors usually has a fourth you have not found.
  • Wrong level throughout: regenerate with the pupil's age, stage and the exact syllabus terms, rather than editing every question.
  • Mostly off-syllabus content: start again and paste the topic list into the prompt.
  • Images carry the learning: remove the images and use your own, or drop the sheet.

An illustrative case: a tutor asked for a sheet on simple interest and got one that switched to compound interest from question five, with a key that mixed the two formulas. Correcting it would have meant rewriting half the questions and all of the key. Regenerating with "simple interest only, formula I = P × R × T ÷ 100, no compound interest" took one minute, and the new sheet needed a single correction.

Pupil details stay out of the prompt

Differentiation tempts tutors to paste in a pupil's name, report comments or learning-support notes so the AI can "tailor" the sheet. It does not need any of that. Describe the need instead: "aged 12, strong at arithmetic, finds long word problems hard, reads slowly, needs larger spacing". That gives the model everything useful and nothing identifying.

Before and after, to show the difference:

  • Before: "Make a worksheet for Aisha, aged 12, who has dyslexia and failed her last test, her mum says she panics with fractions."
  • After: "Make a worksheet for a 12-year-old who reads slowly and loses confidence with fractions. Short sentences, one question per box, start with three easy wins."

The second prompt produces a better sheet and shares nothing about a child. If your tutors use AI for reports as well, the same principle applies there, and writing parent progress reports with AI shows how to keep children's details out of the tool.

Keeping a bank so each sheet is checked once

The real saving comes from never checking the same sheet twice. Store every corrected worksheet with its topic, pupil stage, date checked and who checked it. Over a term, a tutor who makes three sheets a week builds a bank of thirty-plus verified resources, and the next pupil on that topic gets a proven sheet in seconds.

Keep a one-line error log beside the bank: "Q7-type rise/fall questions: key wrong 3 times", "chemistry equations: often unbalanced", "images: labels misspelled". After a month you will know exactly where your assistant fails for your subjects, and you can add a line to your prompt for each one. That is the difference between checking AI worksheets forever and checking them less each month.

Questions tutors ask about checking AI worksheets

Should I tell parents a worksheet was made with AI?

There is no rule that says you must, but honesty helps if a parent asks. A simple line works: you use AI to draft practice material, then check every question and answer yourself before it reaches their child. Parents mostly care that the work is correct and pitched right, so the checking routine is the part worth mentioning.

Is it faster to write worksheets myself than to check AI ones?

For short sheets on topics you teach every week, often yes, because you already have a bank. AI saves time on variety: fresh numbers, new contexts, differentiated versions of the same sheet. If checking a sheet regularly takes longer than fifteen minutes, your prompt is too loose or the topic is too specialised for a general chat assistant.

Can I sell or share worksheets that AI helped me make?

You can usually share them, but check two things first. OpenAI and Anthropic, for example, assign you whatever rights they hold in the output, which may be limited, so you cannot assume full copyright protection. And make sure the sheet does not reproduce past exam questions or textbook passages word for word, since those belong to someone else.

Which AI assistant makes the fewest worksheet errors?

Rankings change with every model update, so test rather than trust a list. Give two assistants the same prompt for a topic you know well, then run the four-pass check on both. Count errors in the answer keys over five or six sheets. The one with fewer key errors on your subjects is the right one for you.

Further reads

Sources: Microsoft Support page on readability and level statistics in Word; OpenAI and Anthropic terms on output ownership; worked maths and science examples checked by hand.

Want a worksheet workflow your tutors can trust?

On a 1:1 call we can map how your tutors produce practice material now, write prompts that give checkable output, and set up a simple sign-off step so no unchecked sheet reaches a pupil.

Book a 1:1 call with me