Ask ChatGPT the same question twice and you can get two different answers — different wording, different sources, sometimes a different set of recommended brands. That's not a malfunction, and it's not you imagining it. The same prompt is almost never the same input twice: the model serving you can change, the live web it reads can change, your own history is quietly folded in, and a layer of genuine randomness sits under all of it. This post walks through the four causes — and why, if you care about whether AI mentions your brand, the variation changes what "checking" even means.
TL;DR
- ChatGPT is not designed to repeat itself. OpenAI's own developer documentation says plainly that "determinism is not guaranteed."
- Four things vary between runs: which model answers you (routing and staged rollouts), what a live web search retrieves, what the engine knows about you (memory, custom instructions, location), and the randomness of sampling itself.
- Even the settings developers use to pin answers down don't fully work — how servers batch requests together changes the arithmetic enough to change the words.
- You can reduce the variation with fresh sessions and specific prompts. You can't remove it.
- For your brand, one answer is a sample, not a verdict. The number that matters is how often you're in the answer across runs.
Is ChatGPT supposed to give the same answer every time?
No — and this isn't a secret. OpenAI's developer documentation on reproducible outputs states that "determinism is not guaranteed," and attributes it to "the inherent non-determinism of our models." That's the documentation for developers who are explicitly asking for repeatable answers and using every control OpenAI provides. The consumer product you type into makes no such attempt.
The system is built for flexible conversation, not for replaying an identical answer on demand. Once you accept that repeatability was never the design goal, the interesting question becomes where the variation comes from — because each source has a different implication for anyone trying to measure what these engines say. There are four.
Which model actually answered you?
"ChatGPT" is not one frozen model. Since GPT-5, a real-time router decides which model handles each message — a fast model for straightforward questions, a slower reasoning model for hard ones — based on the conversation type, its complexity, and even phrases in your prompt. The router itself is continuously retrained, so the dividing line moves.
On top of routing, OpenAI ships model updates in stages and runs experiments on subsets of users. Two people asking the identical question — or you, on Tuesday and again on Friday — can be served by genuinely different models without any visible sign of it. Different model, different answer. This is the variation source you have the least control over and the least visibility into.
Did it search the web this time?
When ChatGPT answers a question with current or commercial stakes — which includes almost every buying question — it often searches the live web and builds the answer from what it retrieves. Two things vary here. First, whether it searches at all can differ between runs of the same prompt. Second, the web itself moves: pages get published, updated, and re-ranked continuously, so the handful of sources retrieved this hour may not match last hour's.
Because the answer is synthesized from those sources, the answer moves with them. A reply grounded in live retrieval is only as stable as the pages it just read. This is also the variation source that should interest you most as a brand, because it's the one you can actually influence: the sources engines keep retrieving for your category's questions are where your visibility is decided.
Why do two people get different answers to the same prompt?
Because they aren't really sending the same prompt. ChatGPT can carry memory between conversations, apply your custom instructions to every reply, and factor in your approximate location. Each question arrives wrapped in a different envelope of context, so the "same" question from two accounts is two different inputs.
For brand checks, this one is a trap with a known fix: your own account has watched you research your own company for months, which can skew answers toward brands you've already discussed. Testing from a fresh or logged-out session gets you closer to what a stranger sees — it's the first rule of the manual method for checking whether AI mentions your brand. Personalization also means your customer in another city or country may get an answer you will never personally observe.
Where does the built-in randomness come from?
Two layers. The first is sampling: language models generate text by choosing each next token from a probability distribution rather than always taking the single most likely one. That's a deliberate design choice — it makes output less repetitive and more natural — and it means some variation is intended behavior, not noise on top of the system.
The second layer is stranger: even when developers pin the sampling down, outputs still differ. Research from Thinking Machines Lab on defeating nondeterminism in LLM inference traced the remaining variation to server batching — your request is processed alongside whatever other requests happen to arrive at the same moment, and the batch's size changes the order of floating-point arithmetic just enough to nudge the numbers, which can change a token, which changes everything after it. Your answer literally depends on who else was asking something at the same time. Making inference fully repeatable required rewriting core GPU operations — which tells you how deep the non-determinism runs in every production system you'll ever query.
What does this mean for checking your brand?
Here's the part most brands miss. If ChatGPT can produce three different brand lists on three runs of the same buying question, then checking once tells you almost nothing. Catch the run where you're named and you'll relax; catch the run where a competitor sweeps the list and you'll panic. Neither reaction is justified, because neither run is "the" answer — there is no single answer to be found.
What exists instead is a distribution. Across many runs of the questions your buyers ask, you're in the answer some fraction of the time — and that fraction is your actual AI visibility. A brand that flickers in and out is usually sitting near the model's evidence threshold: enough corroboration to surface sometimes, not enough to surface reliably. That's a meaningfully different diagnosis from a brand that's absent on every run, which points to one of the five causes of AI invisibility rather than to variance.
The practical consequences:
- Never act on one answer. A single surprising result — good or bad — deserves two or three re-runs before it deserves a reaction.
- Track trend, not snapshots. One good answer is luck. A rising share of answers, checked the same way over months, is progress you can trust.
- Expect the flicker at the margin. If a competitor appears in some runs and you appear in others, you're both near the threshold — and the third-party evidence you earn next is what tips more runs your way.
If you haven't measured your baseline at all yet, start with the five-minute AI visibility audit — it covers how to read each outcome you'll hit.
Can you make ChatGPT give more consistent answers?
You can narrow the variation, not eliminate it. Specific, detailed prompts constrain the model more than vague ones. Fresh sessions with memory and custom instructions off remove the personalization layer. Asking within one conversation holds more context constant than starting over. Do all three and the answers will cluster more tightly — but routing, live retrieval, and sampling remain, so "run it again" can always legitimately produce something different.
For measurement, though, consistency is the wrong goal. Your buyers aren't optimizing their prompts for repeatability — they're asking naturally, on different accounts, on different days, through different models. The variance isn't an obstacle to measuring your visibility; the variance is why the honest measurement is a rate across many natural runs rather than one carefully staged query. The same applies beyond ChatGPT: Perplexity, Gemini, and Google's AI Overviews all generate answers from live sources, and all vary between runs for the same reasons.
Sampling that distribution by hand every month is doable — the manual method linked above is exactly that. Doing it continuously is what a tool is for: ClappX AnswerX runs your category's buying questions across the engines repeatedly, so you see how often you're the answer — not whether you got lucky once.
See how AI describes your brand today.
Free scan of your paid waste and your AI visibility. 60 seconds, no card, no call.
Run free scan →Common questions
Does ChatGPT give everyone the same answer?
No. Different users can be routed to different model versions, trigger different live web searches, and carry different memory, custom instructions, and locations. The same question can return different answers to different people — and to the same person at different times.
Why does ChatGPT give different answers to the same question?
Four reasons: a real-time router and staged rollouts mean a different model can serve each run; live web search retrieves changing sources; personalization folds in your memory, instructions, and location; and sampling plus server batching add genuine randomness. OpenAI's own docs state determinism is not guaranteed.
How do I make ChatGPT's answers more consistent?
You can narrow the variation — specific prompts, fresh sessions with memory and custom instructions off, staying in one conversation — but you can't remove it. Routing, live retrieval, and sampling remain, so re-running the same prompt can always produce a different answer.
Does answer variation happen in Perplexity and Google AI Overviews too?
Yes. Any engine that retrieves live sources and generates its answer varies between runs for the same reasons. That's why AI visibility is measured as how often you appear across many runs, not whether you appeared in one.