It rolls dice
The model does not pick the single best next word; it samples among good ones. Two runs diverge early and end up in different places. Temperature settings control this in the API; in a chat app you do not see them.
Ask ChatGPT the same thing twice and you may get two answers that disagree. That is not a bug in your account; it is how the model works. Each reply is sampled from many plausible continuations, small changes in wording steer it, and everything earlier in the conversation shapes what comes next. This page explains the three causes, what the variation tells you, and how to turn it from a nuisance into a check.
Free to start. Same question, several models, same conversation in MultipleChat.
The model does not pick the single best next word; it samples among good ones. Two runs diverge early and end up in different places. Temperature settings control this in the API; in a chat app you do not see them.
"Is this a good idea?" and "What is wrong with this idea?" are the same question to you and opposite prompts to the model. Word order, examples and even politeness shift the answer.
Everything above in the chat, including a joke or an earlier assumption, shapes what comes next. A fresh chat and a long chat answer differently.
Vendors update models without changing the name. The ChatGPT you used in March may not be the one you use today. If an answer you relied on changes, that can be why.
Name the customer, the period, the format. Vague questions have many plausible answers; precise ones have fewer.
An answer grounded in a document you uploaded varies far less than one from memory. Put the file in a Project.
"Show your working, then the answer" narrows the sampling to answers that follow from the working.
What survives three runs is the stable core; what varies is the uncertain part.
If the same model gives you different answers, the question has uncertainty in it. That is worth knowing. Run it a second time in a fresh chat, then ask a different model cold, and compare all three. What is consistent across runs and models is what you can rely on; what changes is where to look. In MultipleChat this costs three messages, because the context is shared and the model switch is a selector, not a new app. See why models disagree for reading the differences task by task.
The same method works for facts, code, advice and writing. It costs one or two messages in MultipleChat because every model reads the same conversation, files and Project.
Switch the model from the selector and ask again; the history carries over. Pick a model from a different family: ChatGPT then Claude, or Claude then Gemini. Same-family models share blind spots.
"Argue against this answer as strongly as you can" produces more than "do you agree?". Models are agreeable by default; you have to invite disagreement.
Where the two agree, confidence is warranted. Where they differ, that is the claim, the assumption or the line of code to check. Usually it is the one that mattered.
Facts get a source you open (Perplexity cites; Auto Verification on paid plans runs an independent pass). Code gets a test you run. Advice gets a third model or a person.
The question, both answers and the sources stay together, so the next time the topic comes up the checking is already done.
The free plan gives you the major model families and enough messages to make second opinions a habit, which is the whole point of this page. The paid plans add the latest model in each family, which disagree in more useful ways and invent less; Perplexity Sonar Pro for sourced facts; and Auto Verification, an independent verification pass that runs when factual confidence matters, without you orchestrating it.
Newer models hold your whole Project in view instead of the last few messages, reason through trade-offs instead of picking the first plausible answer, and are wrong less often and with less confidence. For the questions on this page, that is the difference between advice that sounds right and advice you can act on.
Pro from $20 a month, cancel any time. The free plan stays free.
Use when consistency matters.
No. Sampling randomness is part of how it generates text. The same is true of Claude, Gemini and every other model.
Not in the chat app. Through the API, a temperature of zero reduces variation but does not eliminate it. Precise questions and source material help more.
Often neither is fully right; the difference marks the uncertain part. Get a source, a second model or a test for that part.
Agreeableness. It tends to defer to you. That is a reason to ask a second model cold rather than argue with the first.
It gives you the tools to find the consistent part: rerun, second model, grounding in your files, all on the same context. Consistency comes from the method, not the app.
The major model families and a daily message allowance. The latest models and Auto Verification are on paid plans.
What survives is the answer. Free to start.
Also: why AI models disagree · is ChatGPT always right? · should I ask two AIs?
Continue learning
Run the same prompt through ChatGPT, Claude, Gemini and Grok before trusting one answer.
Core featureLet several models draft, challenge and verify each other instead of trusting one answer.
ProjectsKeep files, instructions, code and chats together so every model works from the same context.
FeaturesModels, Projects, collaboration, verification, Studios and images in one workspace.
TrustNot on purpose. Confidently wrong, which is worse, and how to catch it.
TrustThe most useful thing they produce, task by task.
TrustSeven warning signs and a checklist from ten seconds to ten minutes.
TrustSplit, source, open, record. Five minutes, claim by claim.
TrustWhen a second opinion is worth it, and how to ask so it does not just agree.
TrustNo ranking survives a month. The setup that beats any single model.
FeatureModels draft, challenge and verify each other.
FeatureAn independent check on factual claims.