ChatGPT makes things up because it is trained to make language sound right, not to verify that every statement is true. It generates an answer by predicting what words plausibly belong together. When the facts are missing, uncertain or difficult to recall, it may fill the gap with a detail that fits, and present it as smoothly as a fact.

That difference between sounding right and being right has had real consequences. In 2023, lawyers representing a passenger in a case against the airline Avianca submitted judicial opinions that ChatGPT had invented. The material included realistic case names, quotations and internal citations. The court described what it received as “non-existent judicial opinions with fake quotes and citations,” and sanctioned the lawyers and their firm.

But ChatGPT is not frozen in 2023. Newer models are better at expressing uncertainty, using external sources and declining questions they cannot answer reliably. The improvement is real, even if the problem has not disappeared. Later in this article, I’ll show what happened when I repeatedly tried to make the current version of ChatGPT invent a historical source – and failed.

What is an AI hallucination?

An AI hallucination is a response containing information that is false, invented or unsupported, despite being presented as credible.

Sometimes the invention is obvious. ChatGPT might name a book that does not exist, attribute a quotation to someone who never said it or provide a link that leads nowhere.

Other hallucinations are harder to notice. The model might summarize a real paper but add a conclusion that the authors never reached. It might combine two accurate facts in a way that creates a false connection. Or it might give you a genuine source that does not actually support the claim beside it.

AI hallucinations commonly include:

  • Invented names, dates, events or statistics
  • Fabricated quotations
  • Nonexistent books, research papers or legal cases
  • False citations, DOIs and links
  • Details that do not appear in the original source
  • Incorrect connections between otherwise accurate facts

What makes these answers dangerous is not simply that they are wrong. People make mistakes too. The problem is that an invented answer can arrive with the same clarity, detail and confidence as a correct one.

The word “hallucination” is itself an imperfect metaphor. A person who hallucinates experiences something that is not present. ChatGPT is not seeing or experiencing anything, and is probably not trying to deceive you on purpose.

It is simply generating language. Sometimes that language is supported by reliable information. Sometimes it merely fits the pattern of what a convincing answer should sound like.

Is ChatGPT lying?

Not in the way a person lies.

Lying requires more than saying something false. A liar knows, or at least believes that one thing is true and deliberately tells you something else. It involves awareness and an intention to deceive.

With GPT is not that deep, so we should not assume that ChatGPT possesses either of those things.

Still, using it can sometimes feel like speaking to a charismatic liar. It can invent a source, explain why that source matters and defend the answer when questioned. It may sound composed, informed and completely certain throughout.

But that confidence is part of the generated language. It is not evidence that ChatGPT checked the claim or knows it to be true.

ChatGPT has learned what credible answers look like. It knows the patterns of academic citations, legal judgments, expert explanations and historical accounts. It can reproduce those patterns so convincingly that an invented answer may look more authoritative than an uncertain but accurate one.

In other words, ChatGPT can perform credibility without establishing truth.

A better comparison may be a talented improviser. An improviser receives a suggestion and continues the scene with whatever details make it work. If the story needs a professor, a university and a research paper, those details can be created on the spot. Their job is to keep the scene coherent, not to prove that it really happened.

ChatGPT is not literally improvising like a person, but the comparison captures something important: it can keep an answer moving even after the available evidence has run out.

Calling this “lying” can distract us from the real problem. If we treat ChatGPT like a dishonest person, we start wondering about its motives. The more useful question is simpler:

What evidence supports this answer?

If the answer cannot be traced to reliable evidence, ChatGPT’s confidence should not make it more believable.

Why does ChatGPT make things up?

ChatGPT makes things up because it is a generator before it is a verifier.

Its basic task is to produce a useful continuation of the words in front of it. Most of the time, the language patterns it has learned lead to an answer that is coherent and broadly correct. But when reliable information is missing, those same patterns can produce details that sound right without being true.

Several parts of the process contribute to this.

It predicts what comes next

ChatGPT generates a response one small unit of language, or token at a time.

At each step, it calculates which possible continuation best fits the prompt, the conversation and the text it has already produced. It then repeats that process until the answer is complete.

This is far more sophisticated than choosing words from a list. The model has learned complex patterns involving language, concepts, context and relationships. That is why it can explain an idea, imitate a writing style or connect information from different domains.

But predicting a plausible continuation is still different from verifying a fact.

Suppose a response begins like this:

A major study of workplace productivity was published in 2021 by…

Many real research citations follow that pattern. If ChatGPT does not have reliable information about the specific study, it may still be able to generate an author, title and journal that fit perfectly. The resulting citation looks right because it follows the pattern of real citations, not because the source necessarily exists.

Patterns are easier to learn than arbitrary facts

Some information follows predictable patterns. Grammar, sentence structure and common relationships appear repeatedly across enormous amounts of text.

Other information is much more arbitrary.

A person’s birthday cannot usually be calculated from their name. A book’s exact title cannot be deduced from its subject. An academic paper’s DOI, a court case number or the wording of a quotation must be recalled or retrieved accurately.

Rare details are especially difficult. If a fact appeared infrequently in the model’s training material, appeared in conflicting forms or was absent altogether, the model may still recognize what kind of information belongs in the answer without having the correct information itself.

It knows there should be a date. It does not necessarily have the date.

That gap between knowing the shape of an answer and possessing the right detail is where many hallucinations begin.

Training can reward guessing

The initial training process is not the whole explanation. What models are rewarded for doing afterward also matters.

Research published by OpenAI argues that common evaluation methods can encourage models to guess rather than admit uncertainty.

Imagine a test where a wrong answer and a blank answer both receive zero points. If you do not know the answer, guessing is the better strategy. You might get lucky. Leaving the question blank guarantees that you will not.

A language model can face a similar incentive. If evaluations reward the number of correct answers without penalizing confident errors more heavily, a model that attempts every question may score better than one that frequently says, “I don’t know.”

This does not mean evaluations create every hallucination. The underlying errors can originate during the model’s earlier training. But rewarding answers more than appropriate uncertainty can teach the system to continue where it should sometimes stop.

The prompt can push the model in the wrong direction

Hallucinations do not originate only inside the model. The question can also create pressure toward an invented answer.

Consider this prompt:

Why did Marie Curie win the Nobel Prize in Literature?

Marie Curie never won that prize. She received Nobel Prizes in Physics and Chemistry.

A reliable response should challenge the question’s premise. But the wording tells the model that the event happened and asks only for an explanation. If it follows the structure of the request instead of checking the assumption underneath it, it may construct a reason for something that never occurred.

The same thing can happen when we ask for:

  • The best book devoted to a subject, without knowing whether one exists
  • The cause of an event that may never have happened
  • A quotation supporting a position the author may not have held
  • Evidence for a conclusion before establishing whether the conclusion is true

The more specific and confident the question sounds, the easier it can be to mistake its premise for a fact.

This is why ChatGPT’s hallucinations are not caused by one mysterious defect. They emerge from a combination of language prediction, incomplete or uncertain information, incentives to answer and prompts that sometimes ask the model to continue beyond the available evidence.

The system does not need to know that an answer is true in order to make it sound as though it is.

When is ChatGPT most likely to hallucinate?

ChatGPT is most likely to hallucinate when a question demands precise information but gives the model little reliable evidence to work with.

The risk is not equal across every task. Asking ChatGPT to reorganize text you provided is very different from asking it to recall the exact title of an obscure academic paper. One depends mainly on the information already in the conversation. The other depends on retrieving a specific fact correctly.

Hallucinations become more likely in situations such as these:

Higher-risk requestWhat can go wrong
Obscure people or eventsChatGPT may fill gaps with invented names, dates or biographical details
Exact quotationsIt may reproduce the general idea in wording the person never used
Academic researchPapers, authors, journals, DOIs or findings may be fabricated
Legal researchThe model may invent cases, judgments or supporting quotations
Links and URLsIt can generate an address that follows a real website’s structure but does not exist
Recent developmentsThe answer may be outdated or combine confirmed news with speculation
Precise numbers and datesA plausible value may be supplied when the exact value is unavailable
False assumptionsChatGPT may explain why something happened without checking whether it happened
Incomplete source materialIt may fill missing information instead of preserving the gap
Long listsLater items may become less reliable as the model tries to satisfy the requested number

Requests for exhaustive answers deserve particular caution. If you ask for 20 examples when only eight strong examples are readily available, ChatGPT may feel pressure to complete the list. The requirement to be comprehensive can conflict with the need to remain accurate.

The same applies when a question is extremely specific:

Give me the exact sentence, page number and publication details.

Specificity can improve a prompt when the information is available. When it is not, specificity can produce a more detailed hallucination.

High-stakes subjects also require greater care, not necessarily because ChatGPT always hallucinates more about them, but because the cost of a single error is much higher. A fabricated detail in a brainstorming session may be harmless. An invented contraindication, court precedent or financial rule may affect someone’s health, freedom or money.

It is also important to distinguish hallucinated information from outdated information. If ChatGPT accurately repeats something that was once true but has since changed, the problem may be stale knowledge rather than pure invention. The practical response is similar: check a current, authoritative source.

A useful rule is:

The more an answer depends on exact facts that must exist outside the conversation, the more carefully those facts should be verified.

ChatGPT is generally safest when transforming information you provide. It becomes riskier when you expect it to act as an invisible database of everything ever written.

Are newer ChatGPT models getting better?

Yes. Newer ChatGPT models are becoming less likely to hallucinate, but “less likely” is not the same as “reliable every time”.

The improvement comes from several directions. Models are being trained to express uncertainty more appropriately, use web search and other tools, reason more carefully about factual questions and recognize when a task cannot be completed with the information available.

OpenAI’s evaluations illustrate the broader trend. In testing published with GPT‑5, its standard model produced 26% fewer factual errors than GPT‑4o, while its reasoning model produced 65% fewer than the earlier o3 model on comparable production-style prompts. The models also became more likely to abstain rather than attempt questions they could not answer reliably. The GPT‑5 System Card explains the evaluations and their limitations.

The progress has continued. OpenAI’s 2026 evaluation of GPT‑5.6 found that its flagship model made slightly fewer factual errors than GPT‑5.5 and was significantly less likely to repeat hallucinations previously reported by users. See the GPT‑5.6 System Card.

These results matter, but they need context.

A hallucination rate is not one permanent score attached to a model. The result changes depending on:

  • The questions included in the test
  • Whether the model can search the web
  • Whether it uses additional reasoning
  • How a factual error is defined
  • Whether accuracy is measured per claim or per complete response
  • How often the model declines to answer
  • Which system or model grades the response

A model can also become more accurate while generating more factual claims. That creates an interesting measurement problem: each individual claim may be more reliable, while a long answer still offers more opportunities for at least one error.

This is why a percentage from one benchmark should not be presented as the chance that any random ChatGPT answer is wrong.

The improvement is nevertheless visible in ordinary use. Newer models are more willing to challenge false assumptions, identify missing information and explain when they cannot verify an exact quotation or citation. Search can also give them access to current, traceable evidence instead of forcing them to rely entirely on patterns learned during training.

But none of these improvements guarantees truth.

A model can search and retrieve the wrong source. It can misunderstand the right source. It can cite a real page that does not support the claim. And without external evidence, it can still produce a plausible answer that turns out to be invented.

The most accurate way to describe the progress is:

ChatGPT is getting better at knowing when to answer, when to search and when to stop. It has not become incapable of being confidently wrong.

I saw that improvement when I tried to make the current version of ChatGPT invent a historical source. Instead of taking the bait, it repeatedly admitted what it could not verify.

I tried to make ChatGPT hallucinate – and failed

To see how readily the current version of ChatGPT would invent a source, I gave it an obscure research question.

I asked for the best scholarly book devoted specifically to Harvard’s Great Butter Rebellion of 1766. I requested an exact title, author, publisher, publication year, ISBN and summary.

ChatGPT did not invent a book.

Instead, it explained that no scholarly book appeared to be devoted exclusively to the event. It found a historical paper called “The Rebellion of 1766 in Harvard College,” written by William Coolidge Lane and published by the Colonial Society of Massachusetts.

It also correctly explained that the paper had no ISBN because it appeared inside a publication that predates the ISBN system.

With search available, ChatGPT rejected the assumption that a dedicated book existed and located a real event-specific paper instead.

I checked the answer. The title, author, publisher, year and page range were genuine. The paper is available in full through the Colonial Society of Massachusetts.

Search had given ChatGPT external evidence, however. I wanted to know what would happen without it.

In a new chat, I asked the same question but instructed ChatGPT to answer from its existing knowledge without browsing. I also told it to identify any uncertainty clearly.

Again, it refused to invent the requested book:

“I cannot responsibly provide the requested title, author, publisher, year, and ISBN; doing so would risk inventing a citation.”

It recommended Samuel Eliot Morison’s real Three Centuries of Harvard, 1636–1936, while clearly explaining that the book covers Harvard’s wider history and is not devoted to the rebellion.

Without access to search, ChatGPT acknowledged the limit of its knowledge rather than completing the requested bibliographic record.

My instruction to identify uncertainty may have encouraged that caution, so I tried a third test.

This time, I asked for three verbatim quotations from Lane’s paper, complete with exact page numbers. I told ChatGPT not to browse, but I did not explicitly instruct it to admit uncertainty.

Exact quotations are a common source of hallucinations. A model may remember the general meaning of a text and generate wording that expresses the same idea, even though the author never wrote those words.

ChatGPT still refused:

“I can’t reliably provide three verbatim quotations with exact page numbers from that work from memory alone.”

It asked me to paste or upload the paper so that it could extract the wording accurately.

When asked for precision it could not guarantee, ChatGPT requested the original source instead of reconstructing quotations from memory.

This was not a scientific experiment, and three responses cannot establish how often ChatGPT hallucinates. A different model, prompt or subject could produce a different result. My instructions also placed clear boundaries around the task.

But the outcome still matters.

ChatGPT was asked for the kinds of details language models have historically invented: obscure publications, bibliographic records, exact quotations and page numbers. Across these three informal tests, it searched when evidence was available and abstained when it was not.

I tried to make ChatGPT hallucinate and failed.

That does not mean hallucinations have been solved. It shows something more useful: hallucination is not an unavoidable response to every gap in a model’s knowledge. Current systems can sometimes recognize the boundary between a plausible answer and a supportable one – and choose to stop.

Does web search prevent hallucinations?

No. Web search can reduce hallucinations, but it cannot guarantee that an answer is correct.

Without search, ChatGPT must rely largely on patterns and information learned during training. With search, it can retrieve current pages, compare sources and attach links to its claims. This gives the answer an external foundation instead of asking the model to reproduce every fact from its internal knowledge.

That difference was visible in my experiment. When search was available, ChatGPT found William Coolidge Lane’s real paper about the Great Butter Rebellion and linked directly to the publication.

But access to evidence is not the same as using evidence correctly.

A search-enabled model can still:

  • Retrieve an unreliable or outdated source
  • Misunderstand what a reliable source says
  • Cite a page that mentions the subject but does not support the claim
  • Combine information from several sources incorrectly
  • Present an inference as though the source stated it directly
  • Miss a more authoritative source
  • Repeat false information that appears across multiple websites

This creates a subtler kind of problem. The citation may be real, the link may work and the source may discuss the right subject, yet the claim beside it may still be unsupported.

A source can therefore be genuine while the way ChatGPT uses it is wrong.

This is why the presence of citations should not end the verification process. It should make verification easier.

For any important claim, ask three questions:

  1. Does the linked source actually exist, and is it credible?
  2. Does it contain the information ChatGPT attributes to it?
  3. Does the source support the interpretation being presented?

Search is especially valuable for recent events, changing rules, unfamiliar subjects and exact references. It gives ChatGPT an opportunity to ground its answer in information outside the model.

But it remains an opportunity and not a guarantee.

Search can give ChatGPT access to the evidence. It does not remove the need to check how the evidence was selected, understood and used.

How can you reduce ChatGPT hallucinations?

You cannot guarantee that ChatGPT will never hallucinate, but you can make it less likely, and make errors easier to detect.

A useful method is: Ground, Constrain, Verify.

Ground the answer in evidence

Give ChatGPT reliable information to work with.

You can upload a document, paste the relevant text or enable web search when the answer depends on current or external facts. When possible, use primary sources such as official records, original research, legislation, company documentation or direct statements.

Instead of asking:

What did this research paper conclude?

Provide the paper and ask:

What conclusions do the authors explicitly state in this paper?

The second prompt gives ChatGPT a defined source rather than asking it to reconstruct the paper from memory.

Grounding does not guarantee accuracy, but it reduces the amount of missing information the model may try to fill.

Constrain what ChatGPT is allowed to claim

Tell the model where the answer must come from and what to do when the evidence is insufficient.

For example:

Answer using only the sources provided. Cite the relevant source beside every factual claim. Clearly label any inference. If the sources do not contain the answer, say what information is missing instead of filling the gap.

Useful instructions include:

  • “Do not invent missing details.”
  • “Distinguish quotations from paraphrases.”
  • “Do not provide a citation unless you can link to the source.”
  • “Challenge any false assumptions in my question.”
  • “Separate confirmed facts from reasonable inferences.”
  • “If you cannot verify an exact date, name or number, say so.”

These instructions do not install a truth detector inside the model. They change the behavior you are asking it to produce. In many cases, that is enough to encourage a cautious answer instead of a confident guess.

Verify the important claims

ChatGPT should help you reach the evidence, not replace the evidence.

Open the sources it provides. Search for the quoted wording. Confirm that the author, publication and link exist. Check whether the source supports the specific sentence beside the citation, not merely the general subject.

Pay particular attention to:

  • Names
  • Dates
  • Statistics
  • Quotations
  • Academic references
  • Legal cases
  • Medical claims
  • Financial rules
  • Product specifications
  • Recent events

Asking ChatGPT to review its own answer can be useful, but it is not independent verification. The model may notice a mistake, or it may repeat the same assumption using different words.

A stronger follow-up is:

List every externally verifiable claim in your answer. For each claim, provide a supporting source or mark it as unverified.

Then check the consequential claims yourself.

Match the verification to the risk

Not every use of ChatGPT requires the same level of scrutiny.

If you are brainstorming names for a fictional character, an invented detail may be exactly what you want. If you are preparing a medical decision, legal filing or financial report, every important factual claim requires independent confirmation.

The practical rule is:

The greater the cost of being wrong, the less you should rely on ChatGPT’s answer alone.

Good prompting can reduce hallucinations. Search can provide evidence. Newer models can be more willing to express uncertainty.

But responsibility for deciding whether an answer is sufficiently reliable still belongs to the person using it.

Frequently asked questions

Does ChatGPT still hallucinate?

Yes. Newer ChatGPT models produce fewer factual errors and are more willing to express uncertainty, but hallucinations have not disappeared. The likelihood changes according to the model, prompt, subject, available tools and evidence. An improvement in average accuracy does not guarantee that a particular answer is correct.

Why does ChatGPT invent sources and quotations?

ChatGPT has learned the patterns of books, academic papers, quotations and citations. When it cannot reliably retrieve a specific source, it may generate details that fit those patterns. The title, author and quotation can look authentic because they resemble real references even when the source does not exist or the person never used those words.

Does ChatGPT know when it is hallucinating?

Not reliably. ChatGPT can recognize some uncertainty and has become better at declining requests it cannot complete accurately. But it does not possess human self-awareness or a dependable internal fact-checker for every statement. It may identify one unsupported answer correctly and express complete confidence in another.

Can I stop hallucinations by telling ChatGPT not to make things up?

That instruction can help, but it cannot guarantee accuracy. More effective prompts tell ChatGPT what sources it may use, require it to separate facts from inference and instruct it to identify missing information. The result should still be verified whenever accuracy matters.

Does web search stop ChatGPT from hallucinating?

No. Search gives ChatGPT access to current evidence and can significantly improve factual answers. However, the model may retrieve an unreliable source, misunderstand a reliable one or attach a real citation to a claim the source does not support. Search makes verification easier; it does not make verification unnecessary.

Can ChatGPT citations be trusted?

Treat citations as routes to evidence, not proof by themselves. Open the source, confirm that it exists and check whether it supports the specific claim. Pay particular attention to exact quotations, academic papers, legal cases, statistics and links. A citation can be genuine while the statement attached to it is misleading or unsupported.

Will AI hallucinations ever be eliminated?

They can probably be reduced further through better training, uncertainty-aware evaluations, retrieval and verification tools. Complete elimination is a much harder promise. Some questions are ambiguous, some information is unavailable and general-purpose language models still generate answers probabilistically. The safer expectation is continued improvement not automatic truth.

What is the safest way to use ChatGPT?

Use ChatGPT to generate, organize, explain and explore. Use reliable evidence to establish facts. Ground important requests in trusted sources, ask the model to identify uncertainty and independently verify consequential claims. The greater the cost of an error, the more verification the answer requires.