Category leaders

Guide

AI Hallucinations: Why Models Make Things Up

A plain-English guide to AI hallucinations: what a hallucination is, measured hallucination rates for Claude, GPT, Gemini and Grok from independent evaluations, why models make things up, which AI hallucinates least, real-world cases and costs, and what actually reduces the problem.

October 10, 2026 · The AI Rankings

Quick answer: An AI hallucination is a fluent, confident answer from an AI model that is false, such as an invented fact, a fake citation or a made-up quote. Every current model still hallucinates: on Artificial Analysis’s AA-Omniscience test of 6,000 hard factual questions, Claude Opus 5.5 gave a wrong answer rather than declining on 59% of the questions it did not get right, against 51% for GPT-6 Astra and 15% for Google’s Gemini 4 Argon, which is not yet publicly available. Models hallucinate because they are trained to produce the most plausible next words, and OpenAI’s research, published in Nature, shows that accuracy-only grading rewards a confident guess over an honest “I don’t know”. The one caveat: the rate depends heavily on the task, so most models score under 15% when summarising a document they are given but above 45% when answering obscure questions from memory.

This page is the explainer: what a hallucination is, how often current models produce them, why they happen and what reduces them. It is not a model ranking. For the overall ranking, see best AI models. For tools built to research with citations, see best AI for research. For monitoring hallucinations in your own application, see best LLM observability tools. For one-line definitions of the surrounding vocabulary, see the AI glossary.


What is an AI hallucination?

An AI hallucination is an output from an AI model that sounds plausible and is stated with confidence but is factually wrong or unsupported by the source the model was given.

Three things make a hallucination different from an ordinary mistake:

Typical examples:

Where the word comes from

“Hallucinate” became the Cambridge Dictionary Word of the Year for 2023, when Cambridge added the sense that an AI hallucinates when it “produces false information”.

The term is contested. Some researchers prefer “confabulation”, the psychiatric term for filling memory gaps with invented detail, which is the word Farquhar and colleagues used in Nature for one subset of hallucinations. Philosophers Michael Townsen Hicks, James Humphries and Joe Slater argued in Ethics and Information Technology that the right word is “bullshit” in Harry Frankfurt’s sense, because the model is indifferent to whether its output is true. “Hallucination” is the term the industry, the labs and the benchmarks use, so this page uses it.

Types of AI hallucination

Researchers split hallucinations by what the output contradicts. The most widely cited taxonomy, from Huang et al.’s survey (published by ACM), has two families.

TypeWhat goes wrongExample
Factuality hallucination: fabricationThe model invents a fact, person, source or event that does not existA citation to a paper that was never published
Factuality hallucination: contradictionThe model states something that contradicts verifiable real-world factSaying the James Webb telescope took the first picture of an exoplanet
Faithfulness hallucination: contextThe output contradicts or adds to a document the model was givenA summary that includes a figure not in the source report
Faithfulness hallucination: instructionThe output ignores or misreports what the user askedTranslating a passage when asked to summarise it
Faithfulness hallucination: logicThe output contradicts itselfCorrect working in a maths answer followed by a different final number

An older split, from Ji et al., calls output that conflicts with the supplied source “intrinsic” and output that cannot be checked against the source “extrinsic”.

The distinction matters for this page because the two families are measured by different benchmarks, and the benchmarks give very different numbers. Factuality is tested by asking questions with known answers. Faithfulness is tested by giving the model a document and checking whether its output sticks to it.


How often do AI models hallucinate?

Every current model tested hallucinates, and how often depends on the task more than on the model. On hard recall questions answered from memory, most frontier models answer wrongly rather than declining more than half the time they do not know. On summarising a supplied document, most widely used models introduce unsupported content in roughly 3% to 15% of summaries.

Hallucination rates on hard factual questions (AA-Omniscience)

AA-Omniscience, from the independent benchmarking firm Artificial Analysis, is the most current public hallucination measurement covering the current flagships. It asks 6,000 questions across 42 topics in six domains: business, humanities and social sciences, science, engineering and maths, health, law, and software engineering.

It reports three numbers:

ModelDeveloperHallucination rate (lower is better)AccuracyOmniscience IndexSource
Gemini 4 Argon (high)Google15%50%42Artificial Analysis
Grok 4.7 (xhigh)xAI29%47%32Artificial Analysis
Claude Sonnet 5.5 (max)Anthropic47%54%Not publishedArtificial Analysis
GPT-6 Astra (max)OpenAI51%63%43Artificial Analysis
Kimi K3Moonshot AI51%46%Data not available on the current indexArtificial Analysis via the Decoder
GPT-6.1 Sol (max)OpenAI54%Not published42Artificial Analysis
Claude Opus 5.5 (max)Anthropic59%66%46 (highest)Artificial Analysis
GPT-5.6 SolOpenAI92%Data not availableData not availableOur GPT-6 Astra page, from Artificial Analysis

Three things stand out:

  1. Low hallucination and high accuracy are different skills. Gemini 4 Argon has the lowest hallucination rate of any model scoring 45 or more on the Artificial Analysis Intelligence Index, but it gets fewer questions right (50%) than Claude Opus 5.5 (66%) or GPT-6 Astra (63%). Argon’s advantage comes from declining when unsure, which Artificial Analysis describes as being “much more likely to acknowledge when it does not know an answer rather than guess incorrectly”.
  2. The best Omniscience Index score belongs to Claude Opus 5.5 (46), despite its 59% hallucination rate, because it answers far more questions correctly than the lower-hallucination models. Claude Sonnet 5.5 hallucinates less than Opus 5.5 (47% against 59%) but also knows less (54% against 66%).
  3. Rates have fallen sharply in a year at OpenAI. GPT-5.6 Sol answered wrongly instead of abstaining 92% of the time on this test; GPT-6 Astra cut that to 51%.

Newer is not always better on this measure. Anthropic’s own Claude Opus 5 System Card states that Opus 5 “hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall”, and Artificial Analysis measured Kimi K3’s 51% as higher than its predecessor’s.

Hallucination rates when summarising documents (Vectara leaderboard)

The Vectara Hallucination Leaderboard measures faithfulness: each model summarises articles, and Vectara’s HHEM-2.3 detection model checks whether each summary adds anything the article does not support. The current version uses a non-public set of more than 7,700 articles of 50 to 24,000 words. It does not yet include Claude 5-series, GPT-6 or Gemini 4 models.

ModelHallucination rate in summaries
Finix S1 32B (Ant Group), lowest on the board1.8%
GPT-5.4 nano3.1%
Gemini 2.5 Flash-Lite3.3%
DeepSeek-V3.26.3%
GPT-5.47.0%
Gemini 2.5 Pro7.0%
GPT-5.59.3%
Gemini 3.1 Pro Preview10.4%
Claude Sonnet 4.610.6%
Claude Opus 4.712.0%
Grok 4.1 Fast Reasoning19.2%

Small models top this board because summarising a supplied text rewards sticking closely to the source, and smaller models paraphrase less. The ranking order here is close to the reverse of AA-Omniscience, which is why no single number answers “which AI hallucinates least”.

Other independent measurements

StudyWhat it testedHeadline result
HalluHard950 multi-turn questions in law, research, medicine and codingThe strongest setup, Claude Opus 4.5 with web search, still hallucinated about 30% of the time
EBU and BBC news integrity studyMore than 3,000 news answers from ChatGPT, Copilot, Gemini and Perplexity, 22 broadcasters, 14 languages45% had at least one significant issue; 20% had major accuracy errors; Gemini worst at 76%
Tow Center, Columbia Journalism Review1,600 requests to 8 AI search engines to identify a quoted articleWrong on more than 60%; Perplexity best at 37% wrong; Grok 3 worst at 94%
Stanford RegLab and HAI (Journal of Empirical Legal Studies)Legal research tools built on retrieval: Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AIEach hallucinated on 17% to 33% of queries
Mount Sinai (Communications Medicine)300 clinical cases containing one planted fake detail, 6 modelsModels repeated or elaborated the fake detail 50% to 82.7% of the time; a safety prompt cut the average from 66% to 44%
Spracklen et al., USENIX Security576,000 code samples from 16 models19.7% of recommended packages did not exist: 5.2% for commercial models, 21.7% for open-source

What the labs say about their own models

Vendor figures use the labs’ own prompts and grading, so they cannot be compared directly with the independent numbers above. They are still the best guide to direction of travel.

Lab and modelVendor claimIndependent figure
OpenAI, GPT-6.1 SolResponses with a factual error on user-flagged hard prompts fell from 11.4% to 7.7% versus GPT-6 Sol (OpenAI)54% AA-Omniscience hallucination rate
OpenAI, GPT-6 AstraAstra “makes substantially fewer factual errors than GPT-5.6 Sol” (OpenAI system card)51%, down from GPT-5.6 Sol’s 92%
Anthropic, Claude Opus 5.5”Our strongest model on most measures of honesty”; in a report-writing test, 16 of 18 Opus 5.5 reports contained no invented figures or quotes (Anthropic)59%, the highest of the current flagships in the table above
Anthropic, Claude Opus 5”Hallucinates factual claims slightly more than Opus 4.8”Rate rose about 14 points to roughly 50%, per our Opus 5 page
Google, Gemini 4 ArgonNo factuality or hallucination figures in Google’s launch announcement15%, the lowest among leading models (Artificial Analysis)

The clearest conflict is Anthropic’s. Anthropic describes Claude Opus 5.5 as its most honest model, and Artificial Analysis measures it hallucinating more often than Claude Sonnet 5.5, GPT-6 Astra or Grok 4.7. Both can be true: Anthropic’s honesty measures cover deception and invented content in long tasks, while AA-Omniscience measures whether a model guesses on short questions it cannot answer.

OpenAI’s GPT-6 Astra system card says its own figures “should not be interpreted as hallucination rates observed in production”, and the same caution applies to every number on this page.


Why do AI models hallucinate?

AI models hallucinate because a large language model generates the most statistically plausible next words, not verified facts, and because the way models are trained and scored rewards a confident guess over admitting uncertainty. Five causes are well documented.

1. Models predict plausible text, not true text

A large language model is trained to predict the next token (a word or part of a word) across trillions of words of text. Nothing in that objective checks whether a sentence is true. A fluent, well-formatted citation is exactly the kind of text the model has learned to produce, whether or not the paper exists.

OpenAI’s researchers show that some errors are unavoidable for facts the model saw only rarely. In their paper, Adam Tauman Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang argue that if 20% of birthday facts appear exactly once in the training data, a base model should be expected to hallucinate on at least 20% of birthday questions. Patterns that repeat, like spelling, are learned reliably. Arbitrary one-off facts, like a private person’s birthday, are not.

2. Training and benchmarks reward guessing

This is the central finding of the same OpenAI research, peer-reviewed and published in Nature under the title “Evaluating large language models for accuracy incentivizes hallucinations”. Most benchmarks score an answer as right or wrong, and “I don’t know” scores the same as a wrong answer. A model that always guesses therefore beats a model that abstains when unsure.

OpenAI’s explainer uses a birthday example: a model asked for someone’s birthday that guesses “September 10” has a 1-in-365 chance of scoring, while “I don’t know” is guaranteed to score zero. Over thousands of questions the guesser wins on accuracy and loses on errors.

OpenAI’s own SimpleQA comparison shows the trade-off:

Metricgpt-5-thinking-miniOpenAI o4-mini
Abstained (“I don’t know”)52%1%
Accurate22%24%
Wrong (hallucinated)26%75%

o4-mini is two points more accurate and hallucinates almost three times as often. On an accuracy-only leaderboard, o4-mini looks better. The researchers’ proposed fix is to “penalize confident errors more than you penalize uncertainty”, which is the approach AA-Omniscience takes.

3. A misfiring “I know this” signal

Anthropic’s interpretability team traced one hallucination mechanism inside a model. In interpretability research, Anthropic found that Claude 3.5 Haiku has a circuit that is on by default and makes the model say it lacks enough information to answer. When the model recognises a name, a “known entity” feature switches that default off. A hallucination happens when the model recognises a name but knows nothing else about it: the “I can’t answer” circuit is suppressed, and the model fills the gap with plausible invention.

This is why questions about obscure people, small companies and niche papers are the highest-risk prompts. OpenAI’s PersonQA test, which asks about real people, shows the same pattern.

4. Reasoning models can make more claims, and more wrong ones

Reasoning models, which generate intermediate steps before answering, do not automatically hallucinate less. OpenAI’s o3 and o4-mini System Card reported these hallucination rates:

ModelPersonQA hallucination ratePersonQA accuracySimpleQA hallucination rate
OpenAI o116%47%44%
OpenAI o333%59%51%
OpenAI o4-mini48%36%79%

OpenAI’s explanation was that “o3 tends to make more claims overall”, so it produced both more correct and more incorrect statements. Later reasoning models improved: OpenAI said GPT-5 with thinking was about 80% less likely to contain a factual error than o3.

5. Long inputs and bad source material

Hallucination rises as the input gets longer. A study of 172 billion tokens of document question-answering found the best model fabricated content in 1.19% of answers at 32,000 tokens of context, that fabrication nearly tripled at 128,000 tokens, and that no model stayed below 10% at 200,000 tokens. The same study found that setting temperature to zero did not consistently reduce fabrication.

Models also repeat errors in their inputs. The Kalai paper notes that training data “inevitably contains errors and half-truths”, and the Mount Sinai study above shows models elaborating on a false detail planted in the prompt.

Is hallucination a bug that will be fixed?

Not entirely. Two theoretical papers argue it cannot be eliminated in principle: Xu, Jain and Kankanhalli’s “Hallucination is Inevitable” concludes “it is impossible to eliminate hallucination in LLMs”, and Banerjee, Agarwal and Singla’s “LLMs Will Always Hallucinate” calls it “an inevitable feature”. The practical goal is a model that knows when it does not know, and declines, rather than a model that is never wrong.


Which AI hallucinates the least?

No single model wins, because the answer depends on whether you need the model to know things or to stick to a document you supply. On the independent data:

If you need…Best on independent dataEvidenceCaveat
The fewest confident wrong answers among leading modelsGemini 4 Argon15% AA-Omniscience hallucination rateRestricted release, not publicly available; answers fewer questions correctly (50%)
The fewest wrong answers among the generally available flagships in our tableGrok 4.729% hallucination rateLower accuracy (47%) than the Claude and GPT flagships
The most right answers overallClaude Opus 5.566% accuracy, highest Omniscience Index (46)59% hallucination rate: it guesses when unsure
A balance of accuracy and restraintGPT-6 Astra63% accuracy, 51% hallucination rate, Index 43Expensive; see our model page for pricing
Faithful summaries of your own documentsSmaller models on the Vectara board1.8% to 3.3% for the top threeNo Claude 5-series, GPT-6 or Gemini 4 models on the board
Research with checkable citationsSearch-grounded toolsClaude Opus 4.5 with web search hallucinated least on HalluHard (about 30%)See best AI for research for tools

For the full model comparison across capability, price and speed, see best AI models. For chat apps, see best AI chatbots.


Real-world AI hallucination cases

Hallucinations have moved from embarrassing screenshots to court sanctions, refunds and withdrawn government policy. The best-documented cases:

What happenedConsequence
Google’s Bard launch demo said the James Webb Space Telescope took the first picture of a planet outside the solar systemAlphabet shares fell 8% that day, wiping out more than $100 billion in market value (Al Jazeera)
Mata v. Avianca: lawyers filed a brief citing cases ChatGPT invented$5,000 sanction (court opinion)
Moffatt v. Air Canada: the airline’s chatbot invented a bereavement refund policyAir Canada ordered to pay C$812.02 (tribunal decision)
Google AI Overviews told users to eat a small rock a day and add glue to pizza, drawn from satire and a Reddit jokeGoogle restricted satire and forum content in Overviews (The Register)
A syndicated summer reading list in the Chicago Sun-Times and Philadelphia Inquirer included at least 10 books that do not exist out of 15The freelance writer’s contract was ended (SAN)
Deloitte’s report for Australia’s Department of Employment and Workplace Relations contained fabricated references and an invented court quoteDeloitte refunded A$97,587, the final instalment of a contract worth about A$440,000 (CFO Dive)
Couvrette v. Wisnovsky (District of Oregon): 15 fake cases and 8 invented quotations across three briefsAbout $110,000 in fees and fines; the judge called the case “a notorious outlier in both degree and volume” (ABA Journal)
South Africa’s draft National AI Policy cited sources that do not existThe minister withdrew the draft 17 days after publication (Rest of World)
EY Canada’s cybersecurity report on loyalty-programme fraud: GPTZero found 16 of its 27 sources were invented, misattributed or brokenReport withdrawn (Computing)
KPMG’s agentic AI report: GPTZero found only 5 of 45 citations accurately pointed to real sourcesKPMG removed the report (GPTZero)

Hallucinations in court

The largest public record of AI hallucinations is legal. Damien Charlotin’s AI Hallucination Cases database at HEC Paris lists more than 2,100 court decisions worldwide involving AI-fabricated content.

BreakdownCases
Filed by people representing themselvesMore than 1,200
Filed by lawyersMore than 850
Involving judgesAbout 30
United StatesNearly 1,500
CanadaMore than 200
AustraliaMore than 100
United KingdomMore than 70

For legal tools and how they compare on verified citations, see best AI for legal.

Hallucinations in research and code

How often people check

Many users do not verify. The KPMG and University of Melbourne global trust study of more than 48,000 people in 47 countries, found that 66% of employees using AI rely on its output without evaluating its accuracy, and 56% have made mistakes in their work because of AI.


How to reduce AI hallucinations

Nothing removes hallucinations completely, but several methods cut them measurably. Grounding the model in real sources, letting it decline, and checking its output work best together.

For everyday users

  1. Turn on web search or use a search-grounded tool. OpenAI reported that GPT-4o with search scored 90% on SimpleQA, against 63% for GPT-4.5 without search, and on HalluHard the lowest hallucination rates all came from models using web search. See best AI search engines.
  2. Give the model the source. Pasting the document or uploading the file turns a recall question into a summarising task, where hallucination rates on the Vectara board are mostly under 15%.
  3. Tell it that “I don’t know” is acceptable. In the Mount Sinai study, a one-line safety instruction cut the average rate at which models elaborated on a false medical detail from 66% to 44%.
  4. Ask for sources, then open them. A citation is only useful if you check it exists and says what the model claims. The Tow Center study found AI search engines misattributed sources on more than 60% of requests.
  5. Be most sceptical about names, numbers and quotes. Anthropic’s research shows obscure names are where the “I know this” signal misfires, and invented figures and quotations are the most common hallucinations in the legal and consulting cases above.
  6. Keep long documents short where you can. Fabrication nearly tripled between 32,000 and 128,000 tokens of input in the document-QA study.
  7. Do not rely on lowering temperature. The same study found temperature zero did not consistently reduce fabrication. Our AI glossary explains what temperature does control.

For teams building on AI models

MethodWhat it doesMeasured effectLimit
Retrieval-augmented generation (RAG)Retrieves relevant documents and passes them to the model with the questionReduces hallucination; legal RAG tools still hallucinated on 17% to 33% of queries in the Stanford studyA model can still misread a correctly retrieved document; see our embedding models guide for the retrieval layer
Web search groundingLets the model look facts up before answeringGPT-4o rose to 90% on SimpleQA with search (OpenAI)Search results can be wrong; models still misattribute sources
Abstention prompting and trainingTells or trains the model to decline when unsuregpt-5-thinking-mini abstained on 52% of SimpleQA and cut errors to 26%, against 75% for o4-miniFewer answers overall
Chain-of-VerificationThe model drafts an answer, writes verification questions, answers them independently, then revisesMeta reported precision on Wikidata list questions rising from 0.17 to 0.36 with Llama 65B (Dhuliawala et al.)More calls per answer; correct answers per query also fell in the same test
Semantic entropySamples several answers and flags questions where their meanings disagreePublished in Nature; detects one subset of hallucinations, arbitrary confabulationsThe authors say it is “several times more computationally costly” than a single answer
Hallucination detectorsA second model checks output against the sourceVectara publishes an open detector, HHEM-2.1-Open, and scores its leaderboard with the commercial HHEM-2.3; AWS says Bedrock Automated Reasoning checks deliver “up to 99% verification accuracy”No independent evaluation of the AWS figure found; Microsoft calls Azure groundedness detection “not a silver bullet”
Self-reporting (“confessions”)OpenAI trains a second, honesty-only output in which the model reports its own shortcuts and errors (OpenAI)OpenAI reports a 4.4% false-negative rateDetects problems after the fact; does not prevent them
Evaluation and monitoringLogs and scores production outputs for unsupported claimsFinds regressions after model or prompt changesRequires a labelled test set; see best LLM observability tools

The common thread in the research is that the model needs both a reliable source and permission to say it does not know. OpenAI’s Nature paper argues the industry also needs benchmarks that penalise confident errors, because developers optimise for whatever the leaderboards reward.


How to spot an AI hallucination

A hallucination looks like a correct answer, so the checks have to be external. Five warning signs:

Tools that judge whether text was written by AI do not detect hallucinations; they estimate authorship, not accuracy. Our AI detectors page covers what those tools do.


Will AI ever stop hallucinating?

Probably not completely, but the rate of confident wrong answers is falling where labs have targeted it. The evidence:

The practical conclusion: treat any AI answer that matters as a draft to verify, and choose tools that show their sources.


Frequently asked questions

What is an AI hallucination?

An AI hallucination is a confident, fluent answer from an AI model that is false or not supported by its source, such as an invented fact, a fabricated citation or a made-up quotation. The term covers both wrong facts stated from memory and summaries that add details the source document does not contain.

Why does AI hallucinate?

AI language models hallucinate because they generate the most plausible next words rather than looking up verified facts, and because training and benchmarks score a confident guess higher than “I don’t know”. OpenAI’s research, published in Nature, shows that accuracy-only grading rewards guessing, and Anthropic’s interpretability work shows a model can suppress its “I can’t answer” response simply because it recognises a name.

How often does ChatGPT hallucinate?

It depends on the model and the task. On Artificial Analysis’s AA-Omniscience test, GPT-6 Astra answered wrongly instead of declining on 51% of the hard questions it did not get right, and GPT-6.1 Sol on 54%. OpenAI reports that GPT-6.1 Sol produced a factual error in 7.7% of responses to user-flagged hard prompts. Using ChatGPT with search on reduces errors substantially. See our ChatGPT page.

Does Claude hallucinate less than ChatGPT?

It depends on the model. On Artificial Analysis’s AA-Omniscience test, Claude Sonnet 5.5 (47%) hallucinates slightly less than GPT-6 Astra (51%) and GPT-6.1 Sol (54%), while Claude Opus 5.5 hallucinates more (59%) but answers more questions correctly (66% against 63% for GPT-6 Astra) and has the highest Omniscience Index score (46). For the wider comparison, see ChatGPT vs Claude.

Which AI model hallucinates the least?

Among leading models, Google’s Gemini 4 Argon has the lowest AA-Omniscience hallucination rate at 15%, but it is in limited release and not publicly available. Among generally available flagships, Grok 4.7 scores lowest at 29%. On document summarisation, small models such as Ant Group’s Finix S1 32B (1.8%) lead the Vectara leaderboard.

Do reasoning models hallucinate more?

Sometimes. OpenAI’s o3 and o4-mini System Card showed o3 hallucinating on 33% of PersonQA questions and o4-mini on 48%, against 16% for the older o1, because the newer models made more claims overall. Later reasoning models improved, and OpenAI said GPT-5 with thinking was about 80% less likely than o3 to contain a factual error.

Does RAG stop hallucinations?

No. Retrieval-augmented generation reduces hallucinations by giving the model relevant documents, but Stanford researchers found commercial legal research tools built on retrieval still hallucinated on 17% to 33% of queries. A model can misread, misquote or go beyond a correctly retrieved document.

Can you prevent AI hallucinations with a prompt?

A prompt can reduce hallucinations but not prevent them. In a Mount Sinai study, adding a one-line safety instruction cut the rate at which models elaborated on false medical details from 66% to 44% on average. Telling the model it may say “I don’t know”, supplying the source text and asking for citations you then check all help.

Does setting temperature to zero stop hallucinations?

No. Temperature controls how varied the output is, not whether it is true. A study of 172 billion tokens of document question-answering found that temperature zero did not consistently reduce fabrication, and that higher temperatures reduced it for most models tested.

What are examples of AI hallucinations?

Well-known examples include lawyers citing cases ChatGPT invented in Mata v. Avianca, Air Canada’s chatbot inventing a refund policy, Google AI Overviews suggesting glue on pizza, a newspaper reading list of books that do not exist, and Deloitte refunding part of an Australian government contract over fabricated references.

How many court cases involve AI hallucinations?

Damien Charlotin’s database lists more than 2,100 court decisions worldwide involving AI-fabricated content, including nearly 1,500 in the United States. People representing themselves account for more than 1,200 entries and lawyers for more than 850. In one Oregon federal case, the court imposed about $110,000 in fees and fines.

Can AI hallucinations be detected automatically?

Partly. Hallucination detectors such as Vectara’s HHEM check whether a summary is supported by its source document, and semantic-entropy methods flag answers that change in meaning when the question is repeated. Neither catches every hallucination, and AI-writing detectors do not check accuracy at all.

Is “hallucination” the right word?

It is the standard industry term, and Cambridge Dictionary made “hallucinate” its Word of the Year for 2023. Some researchers prefer “confabulation”, and philosophers writing in Ethics and Information Technology argued for “bullshit” in Harry Frankfurt’s sense, because the model is indifferent to truth rather than perceiving something that is not there.

Will AI ever stop hallucinating?

Probably not completely. Theoretical work argues hallucination cannot be fully eliminated in language models, but measured rates are falling: OpenAI’s AA-Omniscience hallucination rate fell from 92% for GPT-5.6 Sol to 51% for GPT-6 Astra. The realistic goal is models that decline when unsure.

← All guides