Why AI Gives Different Answers Each Time, and How to Fix It

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Why AI Gives Different Answers Each Time, and How to Fix It.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Why AI Gives Different Answers Each Time, and How to Fix It.

Chat AI picks each word with a controlled element of chance, so the same question can produce different wording and sometimes different conclusions. Differences also come from memory, web search, attached files and model updates. Fix it by giving a fixed prompt template, worked examples, a list to choose from and a set output format, then test the same input several times.

Some variation is useful: two versions of a client email give you a choice. Variation in facts, figures or classifications isn't. One piece of popular advice is now out of date, too. You can't switch the randomness off in chat apps, and on many current models developers can no longer turn it down through the API either. The fixes that still work are about what you put in and what shape you ask to get back.

Follow me on Instagram@sagnikteaches

Five reasons the same question gets a different answer

1. Sampling. For each next word, the model works out how likely hundreds of candidates are, then picks one of the likely ones rather than always the single top choice. That's deliberate: always taking the top word makes writing repetitive and oddly worse. But a slightly different word early on sends the rest of the answer down a different path, and by the end two runs can reach different conclusions.

Connect on LinkedInSagnik Bhattacharya

2. Hidden context. Chat apps quietly add information to your question: memories from past conversations, custom instructions you set months ago, earlier messages in the same chat. Two colleagues asking the identical question can get different answers because their apps know different things about them.

Subscribe on YouTube@codingliquids

Here's how that played out at an illustrative letting agency. Two staff pasted the same tenant email about a broken boiler into ChatGPT and asked for a reply. One got a short, friendly note promising to "arrange a visit as soon as possible". The other got a formal letter promising an engineer "within 24 hours", because months earlier she'd saved a custom instruction saying "we always commit to 24-hour repair visits" while covering a different set of properties. Neither prompt was at fault. Before blaming the model, open each person's settings and read their custom instructions and saved memories.

3. Web search. If the assistant searches the web, it may read different pages each time, or the same pages after they've changed. Ask "what does a competitor charge for a standard service?" on Monday and again on Thursday and you can get two different figures, one from the competitor's current price list and one from a review written three years ago. For anything that depends on a source, ask the assistant to list the pages it used, and only trust a figure that comes from the company's own page.

4. Small differences in the input. A reworded question, an extra line of context or a newer version of the attached spreadsheet can shift the answer more than you'd expect. An online homeware shop found its product-description prompt returned about 120 words for one lamp and 300 for the next. The only difference was the supplier notes: one spec sheet ran to two pages, and the model matched the length of what it was given. Adding "between 90 and 120 words, however long the notes are" to the prompt fixed it.

5. The model changed. Vendors update the models behind their apps every few months, and depending on your settings some apps switch between a quick mode and a slower reasoning mode according to the question. A prompt that worked well in spring can behave differently in autumn with nothing changed on your side.

Here's a realistic case. A mobile dog groomer used one prompt to turn the next day's bookings into reminder texts, and for months it returned one short message per customer. After an update, the reply started with "Here are your reminder messages:" and ended with "Let me know if you'd like a different tone!", and the groomer's texting tool sent both lines to the first and last customers on the list. One line in the template stopped it: "Return only the messages, one per line, with nothing before or after them."

When variation is fine and when it's a problem

TaskIs variation acceptable?Why
Drafting emails, posts, headlinesYesGetting options is the point
Brainstorming ideas or questionsYesDifferent runs surface different ideas
Summarising a meeting or documentIn wording, not in factsActions, dates and names must match every time
Answering client questions from your own policiesIn tone, not in substanceTwo clients must get the same answer
Categorising transactions, emails or ticketsNoThe same input must go to the same place
Extracting figures from invoices or formsNoThere's only one right number

If a task sits in the bottom rows and follows fixed rules, ask whether it needs AI at all. A rule such as "any payment to this supplier goes to this account" gives the same answer every time, for free. AI versus rule-based automation helps you decide which a task needs; often the answer is rules for the predictable cases and AI only for the leftovers.

A meeting summary that moved the deadline

Summaries sit in the middle of that table, and they're where drift in the facts does the most damage, because nobody goes back to the transcript. An illustrative recruitment agency summarised a client call with a one-line prompt:

Summarise this call transcript and list the actions.
[TRANSCRIPT]

Two runs on the same transcript came back like this (illustrative output, trimmed):

RUN 1
Actions:
- Send three shortlisted CVs to the client by Friday
- Client to confirm salary band

RUN 2
Actions:
- Agency to send shortlist early next week
- Confirm salary band and start date

On the call, the client had said "ideally Friday, but Monday's fine". Run 1 hardened a preference into a deadline, run 2 softened it, and run 2 also added a "start date" action nobody had agreed. Neither said who owned the salary action. The fix is to ask for the facts in a fixed shape and forbid interpretation:

From the transcript, list every action in this format:
Action | Owner (agency or client) | Deadline in the speaker's exact words, in quotation marks, or NONE
Only include actions someone explicitly agreed to. Do not infer deadlines.

Rerun three times, every version listed the shortlist action with the owner "agency" and the deadline "ideally Friday, but Monday's fine". That's less tidy than "by Friday" and far more useful, because the consultant reading it knows exactly how much slack there is.

Why "set the temperature to zero" no longer works

Temperature is a setting that controls how adventurous the model's word choices are: low values favour the most likely words, high values allow more variety. For years the standard advice for consistent answers was to set it to zero. That advice has two problems in 2026.

First, chat apps such as ChatGPT, Claude and Gemini have never offered the setting; it only exists for developers using the API. Second, the newest models have largely dropped it. Anthropic's API documentation says temperature is deprecated on Claude 4.7 and later models and only its default value is accepted. Microsoft's documentation for OpenAI's reasoning models lists temperature among the parameters the GPT-5 reasoning models don't support. Reasoning models are steered with an "effort" setting instead, which changes how much the model thinks, not how predictable its wording is.

Even when a zero setting was available, it never guaranteed identical output. So the reliable fixes are the ones below, and they work in chat apps and automations alike.

Fixes that work, from five minutes to a day

  1. Use a fixed template, not a fresh question each time (5 minutes). Same sections, same order, same wording, with blanks for the parts that change. Much of the variation between colleagues is really variation between their prompts.
  2. Give it the list to choose from (10 minutes). For any classification, paste the allowed answers (your account codes, your email categories) and say that nothing outside the list is acceptable.
  3. Add an "UNSURE" option (1 minute). Without one, the model has to pick something, and its guesses are where the inconsistency lives.
  4. Show three to five worked examples (15 minutes). Examples of inputs with the correct answers, including a tricky one, pin down your rules better than any description of them.
  5. Fix the output format (5 minutes). A table with named columns, or one line per item in a set pattern, stops the model inventing a new layout each time and makes answers easy to compare.
  6. Put standing instructions in a shared project (30 minutes). A ChatGPT or Claude Project, or a Gemini Gem (becoming a skill from November 2026), gives everyone the same instructions and reference files. Don't start a custom GPT for this: OpenAI is retiring them, and they stop running on 11 December 2026. How projects and Gems compare helps you pick.
  7. Keep personal memory out of repeat tasks (5 minutes). Run the task inside its project, or in a chat that doesn't use memory: Claude's incognito chats aren't saved to memory, and ChatGPT's Temporary Chat doesn't create memories. Whether ChatGPT remembers conversations covers the settings.
  8. Switch off web search for any task that should use only the material you've provided.
  9. One decision per prompt. Asking for a category, an explanation and a client email in one go multiplies the ways runs can differ. Split them.
  10. For automations, enforce the format in code (about a day of a developer's time). OpenAI's Structured Outputs and Anthropic's strict tool use can make a model's reply match a fixed schema, the exact fields and allowed values you define. Add a simple check afterwards, such as "the account code must exist in our chart", and send anything that fails to a person.

Client answers that must match your policy

Fixes 2, 3 and 5 matter most when AI drafts replies to customers. An illustrative independent bike shop asked its assistant "Can a customer return a sale item?" without attaching its policy. One run said "Yes, within 30 days with a receipt"; another said "Sale items are usually non-refundable". Both sounded sure of themselves. The shop's real policy takes unused sale items back within 14 days. With nothing to go on, the model filled the gap with what shops in general tend to do. The rewritten prompt ties it to the shop's own words:

Answer the customer's question using ONLY the returns policy below.
Quote the sentence from the policy that your answer relies on.
If the policy does not cover the question, reply exactly:
"Good question. I'll check with the team and come back to you today."

Policy:
[PASTE RETURNS POLICY]

Customer question:
[PASTE]

Run five times, the replies were worded differently but every one said 14 days, unused, and quoted the same sentence. That's the split the table asks for: the tone can move, the substance can't. When the policy changes, update the pasted text in the shared project the same day, or the old rule keeps coming back in replies.

A bookkeeping firm's coding prompt, before and after (illustrative)

Say a bookkeeping firm asks AI to suggest categories for a client's bank transactions. The first attempt is a question typed fresh each time.

BEFORE
What categories should these transactions go in?
[pasted list of 40 transactions]

Run three times, the same line, a payment to an office supplier, came back as "Office costs", "Office expenses" and "Stationery". Two of those aren't even accounts in the client's chart. A card payment to a software company was "Software" once and "Subscriptions" twice.

AFTER
You are coding bank transactions for a small business client.

Allowed accounts (use the code exactly; nothing else is allowed):
400 Advertising | 404 Bank fees | 408 Cleaning | 412 Consulting
429 General expenses | 433 Insurance | 461 Printing and stationery
463 IT software and subscriptions | 469 Rent | 489 Telephone and internet

Rules:
- Card payments to software companies go to 463.
- If you are not confident, answer UNSURE. Do not guess.
- Do not calculate or change any amounts.

Examples:
"CARD 14/08 OFFICE SUPPLIER LTD 23.40" -> 461
"DD WORKSPACE RENT AUG 1,250.00" -> 469
"CARD 02/08 TRAIN TICKET 38.10" -> UNSURE (no travel account in list)

Return a table with exactly these columns:
Date | Description | Amount | Account code | Reason (max 8 words)

Transactions:
[PASTE]

The rewritten prompt doesn't make the model cleverer. It removes the choices the firm didn't want it to make: invented account names, guesses on unclear lines and a new layout each time. Store it in the firm's shared prompts so every bookkeeper uses the same version; ChatGPT prompts for bookkeepers has more in the same style. The account codes shown are placeholders; use the client's own chart.

Run a consistency test before you trust it

Before any AI output feeds straight into records or goes to clients without a person reading it, test it the same way every time.

  1. Collect 20 real items where you already know the right answer, including a few awkward ones.
  2. Run each item three times, in fresh chats or through the automation as it will actually run.
  3. For each item, record whether all three runs agreed, and whether they matched the right answer.
  4. Work out two percentages: agreement (items where all three runs matched each other) and accuracy (items where the answer was right).
CONSISTENCY TEST: [TASK]        Date: ____   Prompt version: ____
Item | Right answer | Run 1 | Run 2 | Run 3 | All agree? | Correct?
1    |              |       |       |       |            |
...
20   |              |       |       |       |            |
Agreement: __ / 20      Accuracy: __ / 20      Serious errors: __

My rule of thumb for reading it: at 95% or more on both, with no serious errors, the task can run with a weekly sample check. Between 80% and 95%, let AI suggest and a person confirm each answer. Below 80%, fix the prompt or keep the task manual. Keep the 20 items and rerun the test monthly, and whenever the vendor announces a model change, because a prompt that passed in spring can drift by autumn.

Here is the test filled in by an illustrative landscaping firm that wanted AI to sort incoming emails into Quote request, Existing job, Complaint, Supplier or UNSURE (six of the twenty rows shown):

CONSISTENCY TEST: Enquiry sorting        Prompt version: 2
Item | Right answer  | Run 1         | Run 2         | Run 3         | Agree? | Correct?
1    | Quote request | Quote request | Quote request | Quote request | Y      | Y
4    | Complaint     | Complaint     | Existing job  | Complaint     | N      | N
7    | Supplier      | Supplier      | Supplier      | Supplier      | Y      | Y
12   | Quote request | Existing job  | Existing job  | Existing job  | Y      | N
15   | UNSURE        | UNSURE        | Quote request | UNSURE        | N      | N
19   | Existing job  | Existing job  | Existing job  | Existing job  | Y      | Y
Agreement: 18 / 20 (90%)   Accuracy: 17 / 20 (85%)   Serious errors: 1 (item 4)

At 90% and 85%, with one serious error, this lands in the "AI suggests, a person confirms" band. The two disagreements were the useful part. Item 4 was a complaint written politely ("just wondering when someone might come back about the fence"), and item 15 was a one-line email with no detail at all. The firm added both as worked examples, plus a rule that any mention of a delay or a missed visit counts as a complaint. Item 12, a returning customer asking for a price on a new patio, was harder to spot: all three runs agreed, and all three were wrong. Only the "right answer" column caught it. A rule saying "new work is a quote request, even from an existing customer" fixed that. Version 3 of the prompt scored 20 out of 20 on agreement and 19 out of 20 on accuracy with no serious errors, so it moved to weekly sample checks.

Using disagreement as a free quality check

Variation has one useful side. When the same item gets different answers on different runs, that item is usually genuinely ambiguous: an unclear description, an unusual supplier, a transaction that could sensibly go two ways. So instead of fighting the variation, use it. Run each item twice, accept the answer automatically only when both runs agree, and send every disagreement to a person.

The cost of running everything twice is small at small-business volumes. Through the API, Claude Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens. Say each transaction takes about 800 input tokens (mostly the chart of accounts and examples) and 20 output tokens: 1,000 transactions run twice is 1.6 million input tokens and 40,000 output tokens, roughly $1.80 in all. The person-time you save by only reviewing the disagreements is worth far more. Setting up human review without slowing down shows how to route those items to the right person.

Keep one limit in mind. Two runs agreeing tells you the item is clear to the model, not that the model is right: item 12 in the landscaping test would have passed a two-run check every time. So keep checking a small random sample of the auto-accepted items (ten a week is plenty at these volumes) against what a person would have chosen. If the sample turns up the same wrong answer twice, a rule or example is missing from the prompt, and adding it will do more than any number of extra runs.

Further reads

Sources: Anthropic API documentation (temperature deprecated on Claude 4.7 and later models; strict tool use); Microsoft Learn documentation on OpenAI reasoning models (unsupported parameters); OpenAI Structured Outputs guide; OpenAI help article on Temporary Chat; Claude support article on incognito chats; Anthropic API pricing.

Need an AI task that gives the same answer every time?

On a 1:1 call we'll look at the task you need to be consistent, decide whether it needs AI or a simple rule, and set up the template, test set and review step to keep it reliable.

Book a 1:1 call with me