AI Models & Technology · AI Hallucination & Accuracy
Are Newer AI Models Less Likely to Hallucinate?
Generally yes — newer AI models tend to hallucinate less often than earlier generations, thanks to improved training techniques, better calibration of uncertainty, and tools like retrieval and search, but hallucination has not been fully eliminated and can still occur even in the most current models.
Key takeaways
- AI labs actively measure and work to reduce hallucination rates across model generations, and reported progress has been real on many standard evaluations.
- Improvements come from multiple sources: better and more curated training data, refined training techniques, and increased use of tools like web search and retrieval that ground answers in real sources.
- Newer models are often somewhat better at expressing appropriate uncertainty or declining to answer rather than confidently guessing, though this varies by model and situation.
- Hallucination rates differ by task type — newer models tend to show bigger improvements on some kinds of questions than others.
- No current model, regardless of how recent, has eliminated hallucination entirely, so verification remains important even with the newest available models.
The General Trend
Newer AI models generally do hallucinate less often than earlier ones, as measured across a range of standard evaluations that AI labs use to track factual accuracy over time. This progress reflects genuine improvements: training data curation has gotten more careful, training techniques have been refined specifically with factuality in mind, and many current AI products now pair models with tools like web search or document retrieval that let a model check against real, current sources instead of relying purely on what it absorbed during training. Put together, these changes have measurably reduced how often newer models produce confidently stated but false information compared to earlier generations.
That said, “less likely” is not the same as “unlikely” or “solved.” Hallucination remains a real, documented behavior in even the most current, most capable models available, and the improvement over time has been gradual and uneven rather than a clean, complete fix.
What’s Actually Driving the Improvement
Several distinct threads of progress have contributed to lower hallucination rates in newer models. On the training side, labs have put increasing effort into curating higher-quality training data and developing techniques aimed specifically at improving factual reliability, rather than treating factuality as an incidental byproduct of general capability improvements. Separately, techniques for calibrating a model’s expressed confidence — essentially, training it to better recognize and communicate when it’s on less certain ground, rather than defaulting to confident-sounding answers regardless of actual certainty — have improved in newer model generations, though this remains imperfect.
Perhaps the most practically significant shift, though, has been architectural rather than purely about the model itself: pairing models with tools like live web search and retrieval-augmented generation. When a model can look up and reference an actual current source rather than generating an answer purely from trained-in memory, the opportunity for hallucination on factual questions drops substantially, because the model has something concrete to ground its answer in rather than needing to generate a plausible-sounding fact from scratch.
It’s also worth noting that hallucination reduction hasn’t been uniform across all types of questions. Some categories of tasks — like broad, well-documented general knowledge — have seen bigger measured improvements than others, such as narrow, highly specific factual recall or requests for precise citations, where the underlying challenge of thin or inconsistent training data persists regardless of how advanced the model otherwise is.
What This Means for How You Should Use AI Today
Even accounting for real progress, the sensible practical stance hasn’t changed: treat AI-generated factual claims, especially specific figures, citations, and niche details, as something to verify rather than accept outright, regardless of how new or advanced the model is. The fact that hallucination is less frequent in newer models is genuinely useful and worth knowing, but “less frequent” still means it happens, and there’s no reliable way to tell, just from reading an answer, whether a particular claim is one of the accurate ones or one of the mistakes.
Bottom Line
Newer AI models generally do hallucinate less than earlier generations, thanks to better training, improved calibration, and tools like retrieval and search — but hallucination hasn’t been eliminated, so verifying specific factual claims remains a good practice regardless of how recent or capable the model is.
Look Up AI Terms
Search plain-English definitions of AI and machine learning terms in our free AI Glossary.
Go deeper
Important caveats
- Hallucination rates are typically measured using specific benchmark tests, and real-world performance on a given question can differ from benchmark results.
- A newer model isn't automatically better at every single task or topic compared to an older one; improvements are generally measured in aggregate across many test cases, not guaranteed case by case.
Frequently asked questions
Does a bigger, more advanced AI model always hallucinate less?
Not automatically. While there's a general trend of improvement across model generations and increasing scale, hallucination reduction depends on specific training choices, evaluation focus, and techniques like retrieval — not simply on a model being newer or larger. Some newer models can still hallucinate on specific types of questions despite overall improvements elsewhere.
Do AI companies publish data about hallucination rates?
Many AI companies do publish evaluation results, including performance on hallucination or factuality benchmarks, as part of model release documentation, though methodologies and specific benchmarks vary between companies, which makes direct comparisons across different providers' models imperfect.
Will hallucination ever be completely solved?
It's an active area of ongoing research, and meaningful progress continues, but many researchers view hallucination as a challenge to be substantially reduced and better managed — through techniques like retrieval and improved calibration — rather than something guaranteed to be completely eliminated given how generative language models work.
Related questions
- Why Do AI Models Sometimes Make Up Facts?
- How Can You Fact-Check an AI-Generated Answer?
- Can AI Models Fabricate Fake Citations and Sources?
- Why Shouldn't You Use AI as Your Only Source for Medical or Legal Advice?
- What is a hallucination rate and how do researchers actually measure it?
- Can two different ai models disagree on the same factual question?
Sources
- [1]Research — Anthropic
- [2]OpenAI Research — OpenAI
Written by Editorial Team
Last updated July 25, 2026
Get one well-sourced answer a week
No spam. Unsubscribe anytime.