Guide
AI Hallucinations: Why Models Make Things Up
A plain-English guide to AI hallucinations: what a hallucination is, measured hallucination rates for Claude, GPT, Gemini and Grok from independent evaluations, why models make things up, which AI hallucinates least, real-world cases and costs, and what actually reduces the problem.
Quick answer: An AI hallucination is a fluent, confident answer from an AI model that is false, such as an invented fact, a fake citation or a made-up quote. Every current model still hallucinates: on Artificial Analysis’s AA-Omniscience test of 6,000 hard factual questions, Claude Opus 5.5 gave a wrong answer rather than declining on 59% of the questions it did not get right, against 51% for GPT-6 Astra and 15% for Google’s Gemini 4 Argon, which is not yet publicly available. Models hallucinate because they are trained to produce the most plausible next words, and OpenAI’s research, published in Nature, shows that accuracy-only grading rewards a confident guess over an honest “I don’t know”. The one caveat: the rate depends heavily on the task, so most models score under 15% when summarising a document they are given but above 45% when answering obscure questions from memory.
This page is the explainer: what a hallucination is, how often current models produce them, why they happen and what reduces them. It is not a model ranking. For the overall ranking, see best AI models. For tools built to research with citations, see best AI for research. For monitoring hallucinations in your own application, see best LLM observability tools. For one-line definitions of the surrounding vocabulary, see the AI glossary.
What is an AI hallucination?
An AI hallucination is an output from an AI model that sounds plausible and is stated with confidence but is factually wrong or unsupported by the source the model was given.
Three things make a hallucination different from an ordinary mistake:
- It is fluent. A hallucinated answer reads exactly like a correct one. There is no change in tone, no hesitation and usually no warning.
- It is specific. Hallucinations tend to be precise: a case name with a docket number, a study with authors and a year, a quote in quotation marks. The precision is what makes them convincing.
- It is invented, not looked up. The model generated the text because it was likely, not because it retrieved it from a verified record.
Typical examples:
- A court brief citing cases ChatGPT had invented, which earned two New York lawyers and their firm a $5,000 sanction in Mata v. Avianca.
- An airline chatbot inventing a refund policy, which a Canadian tribunal ordered Air Canada to honour.
- A coding assistant recommending a software package that does not exist. A USENIX Security study found that 19.7% of the packages recommended by 16 code models were fictional.
Where the word comes from
“Hallucinate” became the Cambridge Dictionary Word of the Year for 2023, when Cambridge added the sense that an AI hallucinates when it “produces false information”.
The term is contested. Some researchers prefer “confabulation”, the psychiatric term for filling memory gaps with invented detail, which is the word Farquhar and colleagues used in Nature for one subset of hallucinations. Philosophers Michael Townsen Hicks, James Humphries and Joe Slater argued in Ethics and Information Technology that the right word is “bullshit” in Harry Frankfurt’s sense, because the model is indifferent to whether its output is true. “Hallucination” is the term the industry, the labs and the benchmarks use, so this page uses it.
Types of AI hallucination
Researchers split hallucinations by what the output contradicts. The most widely cited taxonomy, from Huang et al.’s survey (published by ACM), has two families.
| Type | What goes wrong | Example |
|---|---|---|
| Factuality hallucination: fabrication | The model invents a fact, person, source or event that does not exist | A citation to a paper that was never published |
| Factuality hallucination: contradiction | The model states something that contradicts verifiable real-world fact | Saying the James Webb telescope took the first picture of an exoplanet |
| Faithfulness hallucination: context | The output contradicts or adds to a document the model was given | A summary that includes a figure not in the source report |
| Faithfulness hallucination: instruction | The output ignores or misreports what the user asked | Translating a passage when asked to summarise it |
| Faithfulness hallucination: logic | The output contradicts itself | Correct working in a maths answer followed by a different final number |
An older split, from Ji et al., calls output that conflicts with the supplied source “intrinsic” and output that cannot be checked against the source “extrinsic”.
The distinction matters for this page because the two families are measured by different benchmarks, and the benchmarks give very different numbers. Factuality is tested by asking questions with known answers. Faithfulness is tested by giving the model a document and checking whether its output sticks to it.
How often do AI models hallucinate?
Every current model tested hallucinates, and how often depends on the task more than on the model. On hard recall questions answered from memory, most frontier models answer wrongly rather than declining more than half the time they do not know. On summarising a supplied document, most widely used models introduce unsupported content in roughly 3% to 15% of summaries.
Hallucination rates on hard factual questions (AA-Omniscience)
AA-Omniscience, from the independent benchmarking firm Artificial Analysis, is the most current public hallucination measurement covering the current flagships. It asks 6,000 questions across 42 topics in six domains: business, humanities and social sciences, science, engineering and maths, health, law, and software engineering.
It reports three numbers:
- Accuracy is the share of all 6,000 questions answered correctly.
- Hallucination rate is the share of questions the model did not get right on which it gave a wrong answer instead of declining. A model that says “I don’t know” every time it is unsure scores 0%.
- Omniscience Index runs from -100 to 100: a correct answer adds, a wrong answer subtracts, and declining scores zero. A positive score means the model is right more often than it is confidently wrong.
| Model | Developer | Hallucination rate (lower is better) | Accuracy | Omniscience Index | Source |
|---|---|---|---|---|---|
| Gemini 4 Argon (high) | 15% | 50% | 42 | Artificial Analysis | |
| Grok 4.7 (xhigh) | xAI | 29% | 47% | 32 | Artificial Analysis |
| Claude Sonnet 5.5 (max) | Anthropic | 47% | 54% | Not published | Artificial Analysis |
| GPT-6 Astra (max) | OpenAI | 51% | 63% | 43 | Artificial Analysis |
| Kimi K3 | Moonshot AI | 51% | 46% | Data not available on the current index | Artificial Analysis via the Decoder |
| GPT-6.1 Sol (max) | OpenAI | 54% | Not published | 42 | Artificial Analysis |
| Claude Opus 5.5 (max) | Anthropic | 59% | 66% | 46 (highest) | Artificial Analysis |
| GPT-5.6 Sol | OpenAI | 92% | Data not available | Data not available | Our GPT-6 Astra page, from Artificial Analysis |
Three things stand out:
- Low hallucination and high accuracy are different skills. Gemini 4 Argon has the lowest hallucination rate of any model scoring 45 or more on the Artificial Analysis Intelligence Index, but it gets fewer questions right (50%) than Claude Opus 5.5 (66%) or GPT-6 Astra (63%). Argon’s advantage comes from declining when unsure, which Artificial Analysis describes as being “much more likely to acknowledge when it does not know an answer rather than guess incorrectly”.
- The best Omniscience Index score belongs to Claude Opus 5.5 (46), despite its 59% hallucination rate, because it answers far more questions correctly than the lower-hallucination models. Claude Sonnet 5.5 hallucinates less than Opus 5.5 (47% against 59%) but also knows less (54% against 66%).
- Rates have fallen sharply in a year at OpenAI. GPT-5.6 Sol answered wrongly instead of abstaining 92% of the time on this test; GPT-6 Astra cut that to 51%.
Newer is not always better on this measure. Anthropic’s own Claude Opus 5 System Card states that Opus 5 “hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall”, and Artificial Analysis measured Kimi K3’s 51% as higher than its predecessor’s.
Hallucination rates when summarising documents (Vectara leaderboard)
The Vectara Hallucination Leaderboard measures faithfulness: each model summarises articles, and Vectara’s HHEM-2.3 detection model checks whether each summary adds anything the article does not support. The current version uses a non-public set of more than 7,700 articles of 50 to 24,000 words. It does not yet include Claude 5-series, GPT-6 or Gemini 4 models.
| Model | Hallucination rate in summaries |
|---|---|
| Finix S1 32B (Ant Group), lowest on the board | 1.8% |
| GPT-5.4 nano | 3.1% |
| Gemini 2.5 Flash-Lite | 3.3% |
| DeepSeek-V3.2 | 6.3% |
| GPT-5.4 | 7.0% |
| Gemini 2.5 Pro | 7.0% |
| GPT-5.5 | 9.3% |
| Gemini 3.1 Pro Preview | 10.4% |
| Claude Sonnet 4.6 | 10.6% |
| Claude Opus 4.7 | 12.0% |
| Grok 4.1 Fast Reasoning | 19.2% |
Small models top this board because summarising a supplied text rewards sticking closely to the source, and smaller models paraphrase less. The ranking order here is close to the reverse of AA-Omniscience, which is why no single number answers “which AI hallucinates least”.
Other independent measurements
| Study | What it tested | Headline result |
|---|---|---|
| HalluHard | 950 multi-turn questions in law, research, medicine and coding | The strongest setup, Claude Opus 4.5 with web search, still hallucinated about 30% of the time |
| EBU and BBC news integrity study | More than 3,000 news answers from ChatGPT, Copilot, Gemini and Perplexity, 22 broadcasters, 14 languages | 45% had at least one significant issue; 20% had major accuracy errors; Gemini worst at 76% |
| Tow Center, Columbia Journalism Review | 1,600 requests to 8 AI search engines to identify a quoted article | Wrong on more than 60%; Perplexity best at 37% wrong; Grok 3 worst at 94% |
| Stanford RegLab and HAI (Journal of Empirical Legal Studies) | Legal research tools built on retrieval: Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI | Each hallucinated on 17% to 33% of queries |
| Mount Sinai (Communications Medicine) | 300 clinical cases containing one planted fake detail, 6 models | Models repeated or elaborated the fake detail 50% to 82.7% of the time; a safety prompt cut the average from 66% to 44% |
| Spracklen et al., USENIX Security | 576,000 code samples from 16 models | 19.7% of recommended packages did not exist: 5.2% for commercial models, 21.7% for open-source |
What the labs say about their own models
Vendor figures use the labs’ own prompts and grading, so they cannot be compared directly with the independent numbers above. They are still the best guide to direction of travel.
| Lab and model | Vendor claim | Independent figure |
|---|---|---|
| OpenAI, GPT-6.1 Sol | Responses with a factual error on user-flagged hard prompts fell from 11.4% to 7.7% versus GPT-6 Sol (OpenAI) | 54% AA-Omniscience hallucination rate |
| OpenAI, GPT-6 Astra | Astra “makes substantially fewer factual errors than GPT-5.6 Sol” (OpenAI system card) | 51%, down from GPT-5.6 Sol’s 92% |
| Anthropic, Claude Opus 5.5 | ”Our strongest model on most measures of honesty”; in a report-writing test, 16 of 18 Opus 5.5 reports contained no invented figures or quotes (Anthropic) | 59%, the highest of the current flagships in the table above |
| Anthropic, Claude Opus 5 | ”Hallucinates factual claims slightly more than Opus 4.8” | Rate rose about 14 points to roughly 50%, per our Opus 5 page |
| Google, Gemini 4 Argon | No factuality or hallucination figures in Google’s launch announcement | 15%, the lowest among leading models (Artificial Analysis) |
The clearest conflict is Anthropic’s. Anthropic describes Claude Opus 5.5 as its most honest model, and Artificial Analysis measures it hallucinating more often than Claude Sonnet 5.5, GPT-6 Astra or Grok 4.7. Both can be true: Anthropic’s honesty measures cover deception and invented content in long tasks, while AA-Omniscience measures whether a model guesses on short questions it cannot answer.
OpenAI’s GPT-6 Astra system card says its own figures “should not be interpreted as hallucination rates observed in production”, and the same caution applies to every number on this page.
Why do AI models hallucinate?
AI models hallucinate because a large language model generates the most statistically plausible next words, not verified facts, and because the way models are trained and scored rewards a confident guess over admitting uncertainty. Five causes are well documented.
1. Models predict plausible text, not true text
A large language model is trained to predict the next token (a word or part of a word) across trillions of words of text. Nothing in that objective checks whether a sentence is true. A fluent, well-formatted citation is exactly the kind of text the model has learned to produce, whether or not the paper exists.
OpenAI’s researchers show that some errors are unavoidable for facts the model saw only rarely. In their paper, Adam Tauman Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang argue that if 20% of birthday facts appear exactly once in the training data, a base model should be expected to hallucinate on at least 20% of birthday questions. Patterns that repeat, like spelling, are learned reliably. Arbitrary one-off facts, like a private person’s birthday, are not.
2. Training and benchmarks reward guessing
This is the central finding of the same OpenAI research, peer-reviewed and published in Nature under the title “Evaluating large language models for accuracy incentivizes hallucinations”. Most benchmarks score an answer as right or wrong, and “I don’t know” scores the same as a wrong answer. A model that always guesses therefore beats a model that abstains when unsure.
OpenAI’s explainer uses a birthday example: a model asked for someone’s birthday that guesses “September 10” has a 1-in-365 chance of scoring, while “I don’t know” is guaranteed to score zero. Over thousands of questions the guesser wins on accuracy and loses on errors.
OpenAI’s own SimpleQA comparison shows the trade-off:
| Metric | gpt-5-thinking-mini | OpenAI o4-mini |
|---|---|---|
| Abstained (“I don’t know”) | 52% | 1% |
| Accurate | 22% | 24% |
| Wrong (hallucinated) | 26% | 75% |
o4-mini is two points more accurate and hallucinates almost three times as often. On an accuracy-only leaderboard, o4-mini looks better. The researchers’ proposed fix is to “penalize confident errors more than you penalize uncertainty”, which is the approach AA-Omniscience takes.
3. A misfiring “I know this” signal
Anthropic’s interpretability team traced one hallucination mechanism inside a model. In interpretability research, Anthropic found that Claude 3.5 Haiku has a circuit that is on by default and makes the model say it lacks enough information to answer. When the model recognises a name, a “known entity” feature switches that default off. A hallucination happens when the model recognises a name but knows nothing else about it: the “I can’t answer” circuit is suppressed, and the model fills the gap with plausible invention.
This is why questions about obscure people, small companies and niche papers are the highest-risk prompts. OpenAI’s PersonQA test, which asks about real people, shows the same pattern.
4. Reasoning models can make more claims, and more wrong ones
Reasoning models, which generate intermediate steps before answering, do not automatically hallucinate less. OpenAI’s o3 and o4-mini System Card reported these hallucination rates:
| Model | PersonQA hallucination rate | PersonQA accuracy | SimpleQA hallucination rate |
|---|---|---|---|
| OpenAI o1 | 16% | 47% | 44% |
| OpenAI o3 | 33% | 59% | 51% |
| OpenAI o4-mini | 48% | 36% | 79% |
OpenAI’s explanation was that “o3 tends to make more claims overall”, so it produced both more correct and more incorrect statements. Later reasoning models improved: OpenAI said GPT-5 with thinking was about 80% less likely to contain a factual error than o3.
5. Long inputs and bad source material
Hallucination rises as the input gets longer. A study of 172 billion tokens of document question-answering found the best model fabricated content in 1.19% of answers at 32,000 tokens of context, that fabrication nearly tripled at 128,000 tokens, and that no model stayed below 10% at 200,000 tokens. The same study found that setting temperature to zero did not consistently reduce fabrication.
Models also repeat errors in their inputs. The Kalai paper notes that training data “inevitably contains errors and half-truths”, and the Mount Sinai study above shows models elaborating on a false detail planted in the prompt.
Is hallucination a bug that will be fixed?
Not entirely. Two theoretical papers argue it cannot be eliminated in principle: Xu, Jain and Kankanhalli’s “Hallucination is Inevitable” concludes “it is impossible to eliminate hallucination in LLMs”, and Banerjee, Agarwal and Singla’s “LLMs Will Always Hallucinate” calls it “an inevitable feature”. The practical goal is a model that knows when it does not know, and declines, rather than a model that is never wrong.
Which AI hallucinates the least?
No single model wins, because the answer depends on whether you need the model to know things or to stick to a document you supply. On the independent data:
| If you need… | Best on independent data | Evidence | Caveat |
|---|---|---|---|
| The fewest confident wrong answers among leading models | Gemini 4 Argon | 15% AA-Omniscience hallucination rate | Restricted release, not publicly available; answers fewer questions correctly (50%) |
| The fewest wrong answers among the generally available flagships in our table | Grok 4.7 | 29% hallucination rate | Lower accuracy (47%) than the Claude and GPT flagships |
| The most right answers overall | Claude Opus 5.5 | 66% accuracy, highest Omniscience Index (46) | 59% hallucination rate: it guesses when unsure |
| A balance of accuracy and restraint | GPT-6 Astra | 63% accuracy, 51% hallucination rate, Index 43 | Expensive; see our model page for pricing |
| Faithful summaries of your own documents | Smaller models on the Vectara board | 1.8% to 3.3% for the top three | No Claude 5-series, GPT-6 or Gemini 4 models on the board |
| Research with checkable citations | Search-grounded tools | Claude Opus 4.5 with web search hallucinated least on HalluHard (about 30%) | See best AI for research for tools |
- Best generally available model for low hallucination: Grok 4.7, at 29% on AA-Omniscience, against 47% to 59% for the current Claude and GPT flagships. Gemini 4 Argon’s 15% is lower, but Argon is restricted to vetted users and is not a model most people can choose.
- Best for getting the most questions right: Claude Opus 5.5, with the top Omniscience Index score, provided you verify its answers.
- Best for summarising supplied documents: check the Vectara leaderboard for the model you plan to use, since faithfulness rankings differ sharply from recall rankings.
For the full model comparison across capability, price and speed, see best AI models. For chat apps, see best AI chatbots.
Real-world AI hallucination cases
Hallucinations have moved from embarrassing screenshots to court sanctions, refunds and withdrawn government policy. The best-documented cases:
| What happened | Consequence |
|---|---|
| Google’s Bard launch demo said the James Webb Space Telescope took the first picture of a planet outside the solar system | Alphabet shares fell 8% that day, wiping out more than $100 billion in market value (Al Jazeera) |
| Mata v. Avianca: lawyers filed a brief citing cases ChatGPT invented | $5,000 sanction (court opinion) |
| Moffatt v. Air Canada: the airline’s chatbot invented a bereavement refund policy | Air Canada ordered to pay C$812.02 (tribunal decision) |
| Google AI Overviews told users to eat a small rock a day and add glue to pizza, drawn from satire and a Reddit joke | Google restricted satire and forum content in Overviews (The Register) |
| A syndicated summer reading list in the Chicago Sun-Times and Philadelphia Inquirer included at least 10 books that do not exist out of 15 | The freelance writer’s contract was ended (SAN) |
| Deloitte’s report for Australia’s Department of Employment and Workplace Relations contained fabricated references and an invented court quote | Deloitte refunded A$97,587, the final instalment of a contract worth about A$440,000 (CFO Dive) |
| Couvrette v. Wisnovsky (District of Oregon): 15 fake cases and 8 invented quotations across three briefs | About $110,000 in fees and fines; the judge called the case “a notorious outlier in both degree and volume” (ABA Journal) |
| South Africa’s draft National AI Policy cited sources that do not exist | The minister withdrew the draft 17 days after publication (Rest of World) |
| EY Canada’s cybersecurity report on loyalty-programme fraud: GPTZero found 16 of its 27 sources were invented, misattributed or broken | Report withdrawn (Computing) |
| KPMG’s agentic AI report: GPTZero found only 5 of 45 citations accurately pointed to real sources | KPMG removed the report (GPTZero) |
Hallucinations in court
The largest public record of AI hallucinations is legal. Damien Charlotin’s AI Hallucination Cases database at HEC Paris lists more than 2,100 court decisions worldwide involving AI-fabricated content.
| Breakdown | Cases |
|---|---|
| Filed by people representing themselves | More than 1,200 |
| Filed by lawyers | More than 850 |
| Involving judges | About 30 |
| United States | Nearly 1,500 |
| Canada | More than 200 |
| Australia | More than 100 |
| United Kingdom | More than 70 |
For legal tools and how they compare on verified citations, see best AI for legal.
Hallucinations in research and code
- Academic papers. GPTZero found about 100 fabricated citations in 51 to 53 papers accepted at the NeurIPS conference, about 1.1% of accepted papers. Checking citations is covered in best AI citation generators.
- Software packages. Invented package names are a security risk: attackers can register a hallucinated name and wait for developers to install it, a practice known as “slopsquatting”. The USENIX study found 43% of hallucinated package names came back on every one of 10 repeat queries, which makes them predictable. See best AI for coding for coding tools.
- Transcription. OpenAI’s Whisper speech model invented phrases in about 1% of 13,140 audio segments in a Cornell-led study, and 38% of those inventions included explicit harms such as violence or false medical information. See best AI for transcription and best AI medical scribes.
How often people check
Many users do not verify. The KPMG and University of Melbourne global trust study of more than 48,000 people in 47 countries, found that 66% of employees using AI rely on its output without evaluating its accuracy, and 56% have made mistakes in their work because of AI.
How to reduce AI hallucinations
Nothing removes hallucinations completely, but several methods cut them measurably. Grounding the model in real sources, letting it decline, and checking its output work best together.
For everyday users
- Turn on web search or use a search-grounded tool. OpenAI reported that GPT-4o with search scored 90% on SimpleQA, against 63% for GPT-4.5 without search, and on HalluHard the lowest hallucination rates all came from models using web search. See best AI search engines.
- Give the model the source. Pasting the document or uploading the file turns a recall question into a summarising task, where hallucination rates on the Vectara board are mostly under 15%.
- Tell it that “I don’t know” is acceptable. In the Mount Sinai study, a one-line safety instruction cut the average rate at which models elaborated on a false medical detail from 66% to 44%.
- Ask for sources, then open them. A citation is only useful if you check it exists and says what the model claims. The Tow Center study found AI search engines misattributed sources on more than 60% of requests.
- Be most sceptical about names, numbers and quotes. Anthropic’s research shows obscure names are where the “I know this” signal misfires, and invented figures and quotations are the most common hallucinations in the legal and consulting cases above.
- Keep long documents short where you can. Fabrication nearly tripled between 32,000 and 128,000 tokens of input in the document-QA study.
- Do not rely on lowering temperature. The same study found temperature zero did not consistently reduce fabrication. Our AI glossary explains what temperature does control.
For teams building on AI models
| Method | What it does | Measured effect | Limit |
|---|---|---|---|
| Retrieval-augmented generation (RAG) | Retrieves relevant documents and passes them to the model with the question | Reduces hallucination; legal RAG tools still hallucinated on 17% to 33% of queries in the Stanford study | A model can still misread a correctly retrieved document; see our embedding models guide for the retrieval layer |
| Web search grounding | Lets the model look facts up before answering | GPT-4o rose to 90% on SimpleQA with search (OpenAI) | Search results can be wrong; models still misattribute sources |
| Abstention prompting and training | Tells or trains the model to decline when unsure | gpt-5-thinking-mini abstained on 52% of SimpleQA and cut errors to 26%, against 75% for o4-mini | Fewer answers overall |
| Chain-of-Verification | The model drafts an answer, writes verification questions, answers them independently, then revises | Meta reported precision on Wikidata list questions rising from 0.17 to 0.36 with Llama 65B (Dhuliawala et al.) | More calls per answer; correct answers per query also fell in the same test |
| Semantic entropy | Samples several answers and flags questions where their meanings disagree | Published in Nature; detects one subset of hallucinations, arbitrary confabulations | The authors say it is “several times more computationally costly” than a single answer |
| Hallucination detectors | A second model checks output against the source | Vectara publishes an open detector, HHEM-2.1-Open, and scores its leaderboard with the commercial HHEM-2.3; AWS says Bedrock Automated Reasoning checks deliver “up to 99% verification accuracy” | No independent evaluation of the AWS figure found; Microsoft calls Azure groundedness detection “not a silver bullet” |
| Self-reporting (“confessions”) | OpenAI trains a second, honesty-only output in which the model reports its own shortcuts and errors (OpenAI) | OpenAI reports a 4.4% false-negative rate | Detects problems after the fact; does not prevent them |
| Evaluation and monitoring | Logs and scores production outputs for unsupported claims | Finds regressions after model or prompt changes | Requires a labelled test set; see best LLM observability tools |
The common thread in the research is that the model needs both a reliable source and permission to say it does not know. OpenAI’s Nature paper argues the industry also needs benchmarks that penalise confident errors, because developers optimise for whatever the leaderboards reward.
How to spot an AI hallucination
A hallucination looks like a correct answer, so the checks have to be external. Five warning signs:
- Very specific detail you did not ask for, such as page numbers, docket numbers, exact percentages or direct quotations.
- A source you cannot find. Search for the exact title of any paper, case or book. If it does not appear on the publisher’s site, a court database or Google Scholar, treat it as invented.
- A link that does not open or opens to something else. The Tow Center found AI search tools frequently linked to the wrong article or a broken URL.
- A different answer when you ask again. Semantic-entropy research rests on this: when the same question produces answers with different meanings, the model is likely guessing.
- A confident answer about an obscure person, small company or recent event, especially one after the model’s training cutoff.
Tools that judge whether text was written by AI do not detect hallucinations; they estimate authorship, not accuracy. Our AI detectors page covers what those tools do.
Will AI ever stop hallucinating?
Probably not completely, but the rate of confident wrong answers is falling where labs have targeted it. The evidence:
- Measured rates dropped sharply within one model generation. OpenAI’s hallucination rate on AA-Omniscience fell from 92% for GPT-5.6 Sol to 51% for GPT-6 Astra, and Artificial Analysis measured the cheaper GPT-6 Sol at 60% at maximum effort, per our GPT-6 Sol page.
- Abstention is now a design goal. Gemini 4 Argon’s 15% rate shows a frontier-scale model can decline rather than guess most of the time, at some cost to accuracy.
- The incentive problem has peer-reviewed backing. The Nature publication of OpenAI’s argument puts pressure on benchmark makers to stop scoring “I don’t know” as a wrong answer.
- Progress is uneven. Claude Opus 5 hallucinated more than Claude Opus 4.8, and Kimi K3 more than its predecessor, both while getting more questions right.
- The theoretical floor is above zero. The inevitability papers and the Kalai paper’s singleton argument both imply some error rate on rarely seen facts.
The practical conclusion: treat any AI answer that matters as a draft to verify, and choose tools that show their sources.
Frequently asked questions
What is an AI hallucination?
An AI hallucination is a confident, fluent answer from an AI model that is false or not supported by its source, such as an invented fact, a fabricated citation or a made-up quotation. The term covers both wrong facts stated from memory and summaries that add details the source document does not contain.
Why does AI hallucinate?
AI language models hallucinate because they generate the most plausible next words rather than looking up verified facts, and because training and benchmarks score a confident guess higher than “I don’t know”. OpenAI’s research, published in Nature, shows that accuracy-only grading rewards guessing, and Anthropic’s interpretability work shows a model can suppress its “I can’t answer” response simply because it recognises a name.
How often does ChatGPT hallucinate?
It depends on the model and the task. On Artificial Analysis’s AA-Omniscience test, GPT-6 Astra answered wrongly instead of declining on 51% of the hard questions it did not get right, and GPT-6.1 Sol on 54%. OpenAI reports that GPT-6.1 Sol produced a factual error in 7.7% of responses to user-flagged hard prompts. Using ChatGPT with search on reduces errors substantially. See our ChatGPT page.
Does Claude hallucinate less than ChatGPT?
It depends on the model. On Artificial Analysis’s AA-Omniscience test, Claude Sonnet 5.5 (47%) hallucinates slightly less than GPT-6 Astra (51%) and GPT-6.1 Sol (54%), while Claude Opus 5.5 hallucinates more (59%) but answers more questions correctly (66% against 63% for GPT-6 Astra) and has the highest Omniscience Index score (46). For the wider comparison, see ChatGPT vs Claude.
Which AI model hallucinates the least?
Among leading models, Google’s Gemini 4 Argon has the lowest AA-Omniscience hallucination rate at 15%, but it is in limited release and not publicly available. Among generally available flagships, Grok 4.7 scores lowest at 29%. On document summarisation, small models such as Ant Group’s Finix S1 32B (1.8%) lead the Vectara leaderboard.
Do reasoning models hallucinate more?
Sometimes. OpenAI’s o3 and o4-mini System Card showed o3 hallucinating on 33% of PersonQA questions and o4-mini on 48%, against 16% for the older o1, because the newer models made more claims overall. Later reasoning models improved, and OpenAI said GPT-5 with thinking was about 80% less likely than o3 to contain a factual error.
Does RAG stop hallucinations?
No. Retrieval-augmented generation reduces hallucinations by giving the model relevant documents, but Stanford researchers found commercial legal research tools built on retrieval still hallucinated on 17% to 33% of queries. A model can misread, misquote or go beyond a correctly retrieved document.
Can you prevent AI hallucinations with a prompt?
A prompt can reduce hallucinations but not prevent them. In a Mount Sinai study, adding a one-line safety instruction cut the rate at which models elaborated on false medical details from 66% to 44% on average. Telling the model it may say “I don’t know”, supplying the source text and asking for citations you then check all help.
Does setting temperature to zero stop hallucinations?
No. Temperature controls how varied the output is, not whether it is true. A study of 172 billion tokens of document question-answering found that temperature zero did not consistently reduce fabrication, and that higher temperatures reduced it for most models tested.
What are examples of AI hallucinations?
Well-known examples include lawyers citing cases ChatGPT invented in Mata v. Avianca, Air Canada’s chatbot inventing a refund policy, Google AI Overviews suggesting glue on pizza, a newspaper reading list of books that do not exist, and Deloitte refunding part of an Australian government contract over fabricated references.
How many court cases involve AI hallucinations?
Damien Charlotin’s database lists more than 2,100 court decisions worldwide involving AI-fabricated content, including nearly 1,500 in the United States. People representing themselves account for more than 1,200 entries and lawyers for more than 850. In one Oregon federal case, the court imposed about $110,000 in fees and fines.
Can AI hallucinations be detected automatically?
Partly. Hallucination detectors such as Vectara’s HHEM check whether a summary is supported by its source document, and semantic-entropy methods flag answers that change in meaning when the question is repeated. Neither catches every hallucination, and AI-writing detectors do not check accuracy at all.
Is “hallucination” the right word?
It is the standard industry term, and Cambridge Dictionary made “hallucinate” its Word of the Year for 2023. Some researchers prefer “confabulation”, and philosophers writing in Ethics and Information Technology argued for “bullshit” in Harry Frankfurt’s sense, because the model is indifferent to truth rather than perceiving something that is not there.
Will AI ever stop hallucinating?
Probably not completely. Theoretical work argues hallucination cannot be fully eliminated in language models, but measured rates are falling: OpenAI’s AA-Omniscience hallucination rate fell from 92% for GPT-5.6 Sol to 51% for GPT-6 Astra. The realistic goal is models that decline when unsure.