Why Does ChatGPT Give Different Answers to the Same Question?

Why Does ChatGPT Give Different Answers to the Same Question? — The AI Cheat Sheet

Short answer: because ChatGPT does not look up a stored answer and hand it back. It builds a reply one word at a time, and at every step it picks from a short list of plausible next words rather than always taking the top one. That pick involves a bit of deliberate randomness, so two runs of the same question take slightly different paths. Add saved memory, whatever else is in the chat, and quiet updates to the product itself, and the wonder is that answers repeat as closely as they do.

I asked the same question twice and watched what moved

To see how far the wobble goes, I gave ChatGPT a small task with rules I could check. In one chat, with the default model selected, I asked for three low-cost ways a small neighborhood bakery could pull in more weekday morning customers, as exactly three numbered items, each under 15 words, no introduction. Then I opened a brand-new chat and pasted the identical prompt.

Both replies followed the rules. Neither one used the same sentence twice.

The two runs, side by side

Same prompt, one minute apart

Two ChatGPT replies to an identical bakery prompt, showing the same three ideas worded differently
The same prompt, run twice in two new chats. ChatGPT, default model selected, September 2026.

The structure held perfectly and the three ideas landed in the same order both times. What moved was the substance inside them. Run one aimed the bundle at "the morning commute"; run two aimed it at "the first two hours" — related, but not the same instruction to a shop owner deciding when to staff the counter. The loyalty idea turned into a punch card. Neither run broke a single rule I set, and neither one mentioned that a second run would say something different. That is the part worth sitting with: the formatting you can see stayed put, and the advice you would actually act on quietly changed.

Why it happens, in one plain sentence

A language model works out which words could plausibly come next, then chooses one — and by default that choice is a weighted roll rather than always grabbing the single likeliest option. Take a different word early on and the rest of the sentence has to follow it. "Offer a weekday coffee-and-pastry bundle during the…" can finish several ways, and once one of them is on the page the model commits to it.

This is a design choice, not a defect. Always taking the top-ranked word produces text that is repetitive and oddly flat. The randomness is what makes the writing readable. It just also means you are getting one sample of a good answer, not the good answer.

Four other things that shift the reply

Saved memory and custom instructions

If the tool has stored preferences about you, they are folded into every reply, and they can be edited or cleared without you noticing anything except that the tone changed. This is the same machinery behind the opposite complaint — the one where it forgets a detail you are sure you gave it. I went through what is actually kept and what is not in why ChatGPT forgets what you told it.

Everything earlier in the same chat

A reply in message twenty is shaped by messages one through nineteen. Ask the same question in a long chat and in a fresh one and you should expect different answers, because they are not really the same question.

Whether it looked anything up

These tools sometimes search the web mid-answer and sometimes do not. A run that searched is working from pages that may have changed since the last run; a run that did not is working from training alone. Same prompt, two different sources of truth.

The product moving under your feet

Models get updated, routing between them changes, and the version answering you in March may not be the one answering in June. Nobody sends you a note about it. If an answer you relied on last quarter now reads differently, this is usually why.

When the variation is fine, and when it is a problem

For anything creative or exploratory — subject lines, taglines, names, ways to open a difficult email, angles on a blog post — the variation is the feature. Re-running a prompt is the cheapest way to get a second opinion, and I do it constantly.

It becomes a problem the moment you treat a reply as a lookup. Dates, dosages, prices, legal thresholds, a figure out of a report, the spelling of a name, whether a policy applies to you: if a question has one right answer, the fact that you got a fluent reply tells you nothing about whether it is the right one, and a second run that says something else is the tool telling you so. My two-minute fact-check routine covers what to do about it without turning every answer into a research project.

There is a middle case that catches people out: you got a great answer last week, you did not save it, and you are trying to conjure it back by retyping roughly the same question. That almost never works. Copy good output somewhere permanent the moment you get it.

How to get answers that hold still

You cannot switch the randomness off in the consumer apps, but you can shrink how much room it has to move.

  • Pin the format. My test showed the shape of the answer obeying the rules exactly while the content drifted. Numbered items, word caps, "no introduction", a required heading — all of that sticks. Use it.
  • Write down the constraint you are holding in your head. If "commute hours" was the bit that mattered, that belongs in the prompt. Anything you leave unspecified is something the model gets to re-decide each run.
  • Ask for the variations up front. "Give me eight options" in one reply beats running the same prompt eight times: you get the range in one place, and you can compare them properly.
  • Reuse the prompt verbatim. Keep the ones that work in a note and paste them, rather than retyping from memory. Small rewordings are a bigger source of drift than the randomness is.
  • Keep one chat per job. Consistency within a thread is much better than consistency across new chats, because the context is the same.
  • Ask it to show its working. Requesting the steps or the source alongside the answer gives you something to check, and makes it obvious when two runs disagree on the reasoning rather than just the phrasing.

If you want to go further on the prompt side, the framework in how to write better AI prompts is mostly about removing exactly this kind of ambiguity.

A quick sanity check you can run on anything important

Ask your question, then open a new chat and ask it again, cold. If the two answers agree on the facts, you have some reason to trust them. If they disagree, you have learned something more useful than either answer: this is a question the tool is guessing at, and you need a real source. It takes about thirty seconds and it has saved me from passing on confident nonsense more than once.

What I would change first

Stop thinking of the chat box as a search bar with better manners. A search engine hands you the same ten results twice in a row; this hands you a fresh piece of writing every time, built to sound finished whether or not it is right. Once that lands, the habits follow on their own — save the good outputs, pin the formats, put your real constraints in writing, and run the important questions twice before you act on them.

Related reading

Written by Mitch, a software analyst who tests software for a living. This guide comes from actually using the tool on a real task. Tested on ChatGPT Plus, Claude Max and Gemini (free). More about this site · Corrections: [email protected].

Get the next guide by email

A new tested how-to when there is one — usually a couple a month. No hype, unsubscribe any time.

© 2026 The AI Cheat Sheet  ·  About  ·  Contact  ·  Privacy Policy  ·  RSS
Tested in real accounts. No affiliate links.

Popular posts from this blog

Is It Safe to Paste Work Documents Into ChatGPT?

How to Use AI to Write and Tailor Your Resume and Cover Letter

How to Spot AI Scams and Deepfakes Before They Cost You