How to know if AI is wrong

How to know if an AI is wrong when the answer looks perfectly fine.

You usually cannot tell from the answer. Wrong AI answers are written in the same confident, well-structured prose as right ones; that is the whole problem. What you can do is stop judging the answer and start judging the evidence: does it have a source you can open, a test you can run, or a second model that agrees for its own reasons? This page is the checklist, ordered from cheapest to most thorough, and the signs that should make you slow down.

Free to start. Auto Verification on paid plans runs the check for you.

Warning signs

Seven things that should make you slow down

SignWhy it matters
A very specific number, name or citation you did not give itSpecifics are where confabulation lives. Ask for the source.
No hedging at all on a question that has real uncertaintyReal experts hedge. A model that does not is filling the gap smoothly.
It agrees with you enthusiasticallyAgreeableness is trained in. Ask it to argue the other side.
The answer gets vaguer towards the endIt ran out of knowledge and kept writing.
A library, API or product feature you have never heard ofIt may have invented it, or it may be renamed or removed. Check the docs.
Arithmetic inside proseRecompute. Models reason about numbers as text.
Anything that could have changed recentlyTraining cut-offs. Use a model that searches, and check the date.
The checklist

From cheapest to most thorough

  1. Ask for certainty marks (10 seconds)

    "Mark each sentence as certain, likely or guessed." The model's own guesses are often honest when asked directly.

  2. Ask it to attack its answer (30 seconds)

    "Argue against your previous answer." The counter-argument surfaces the assumption.

  3. Ask a second model cold (1 minute)

    Different family, fresh chat, same Project. Read where they differ. This catches most of what the first two miss.

  4. Get a source or a test (2 to 10 minutes)

    Facts: a source you open, ideally from a search-native model with citations. Code: the failing test. Numbers: recompute in a spreadsheet.

  5. Run an independent verification pass (paid plans)

    Auto Verification in MultipleChat checks factual claims independently when confidence matters. Use it on anything going into a plan, a pitch or a publication.

  6. Ask a person (when it decides money, law or health)

    AI narrows the question; a professional answers it. The first five steps make that conversation shorter and cheaper.

What counts as evidence

Different kinds of answer need different proof

Facts

A source you opened

Not a citation; an opened page that says what the model claims. Perplexity supplies them; you read them.

Code

A test that runs

Especially the case nobody mentioned: retry, empty input, other time zone, huge input.

Advice

Two models and the counter-argument

There is no source for "should I". Two independent judgements plus the strongest case against are the best you can get.

Writing

A reader who is not you

A second model reading as the sceptical audience, then a real person.

Free is the start. The latest models are the difference.

Checking answers is where the latest models earn their price.

The free plan gives you the major model families and enough messages to make second opinions a habit, which is the whole point of this page. The paid plans add the latest model in each family, which disagree in more useful ways and invent less; Perplexity Sonar Pro for sourced facts; and Auto Verification, an independent verification pass that runs when factual confidence matters, without you orchestrating it.

Newer models hold your whole Project in view instead of the last few messages, reason through trade-offs instead of picking the first plausible answer, and are wrong less often and with less confidence. For the questions on this page, that is the difference between advice that sounds right and advice you can act on.

Pro from $20 a month, cancel any time. The free plan stays free.

Free plan
  • Your first Project and file uploads
  • The major AI families to try the workflow
  • A small daily message allowance
Paid plans
  • The latest ChatGPT, Claude, Gemini, Grok and Perplexity models, all on the same Project
  • Far more messages a day, so a working session does not stop halfway
  • Modes where several models draft, challenge and verify each other's answers
  • Auto Verification: an independent check on factual claims
  • Perplexity Sonar Pro for sourced, current facts
  • Collaboration modes where models draft, challenge and verify each other
Start here

Messages for checking an answer

Cheapest first.

CertaintyMark each sentence of your last answer as certain, likely or guessed.
AttackArgue against your previous answer as strongly as you can.
Cold checkAnswer this independently, without assuming any earlier answer is right: [question]
SourcesGive me a source I can open for each claim. Mark any you cannot source.
Failing caseWhat input would make this code fail? Show the case.
ChangedWhich parts of this could have changed since your training? How would I check today?
Honest answers

What people ask

Can an AI tell me when it is wrong?

Sometimes, if you ask directly for certainty marks. It cannot do it reliably on its own; the confident tone is the default.

Is there a tool that detects AI errors?

No tool reads an answer and flags the wrong sentences reliably. What works is independence: a different model, a source, a test. Auto Verification in MultipleChat automates the source check for factual claims on paid plans.

How much checking is enough?

Match it to the cost of being wrong. A draft email: none. A number in a plan: a source. Code going live: a review and a test. Medical, legal or financial: a professional.

Why not just use a better model?

Better models are wrong less often and still wrong invisibly. The method is what protects you; the model changes how often you need it.

Does MultipleChat make this faster?

Yes. Every model reads the same conversation and Project, so the second opinion, the source and the verification pass are each one message rather than a new app.

What is free?

The major model families and a daily message allowance. Auto Verification, Sonar Pro and the latest models are on paid plans.

Stop judging the answer. Judge the evidence.

Certainty marks, a counter-argument, a second model, a source. Two minutes in MultipleChat. Free to start.

Start checking

Continue learning

Open MultipleChat