Ask ChatGPT, Claude, Gemini, and Grok the same hard factual question and you will get four different answers with four different levels of confidence, and at least one of them will probably be wrong. That is not a guess. It is what Artificial Analysis found when it ran its AA-Omniscience benchmark against the current crop of flagship chatbots in 2026, and the spread between the best and worst performer comes out to roughly 28 percentage points. For a reader trying to decide which AI chatbot to trust with research, customer support, or anything resembling real-world advice, that gap is the whole story.
This comparison lines up OpenAI’s GPT-6 Astra, Anthropic’s Claude Opus 4.8 and Claude Fable 5.1, Google’s Gemini 3.8 Flash, and xAI’s Grok 4.5 against every hallucination and factual-accuracy benchmark with public September 2026 data: Artificial Analysis’s AA-Omniscience, the Vectara Hallucination Leaderboard, Google DeepMind’s FACTS Grounding suite, and LMArena’s factuality Elo scores. It also walks through five real incidents, from lawsuits to a benchmark regression Anthropic had to publicly explain, that show what a hallucination actually costs when it reaches a real user. Pricing, specs, use-case picks, and a migration plan for reducing hallucination risk follow at the end.
Don't miss new tech stories on Google
Add Tech Insider once in the Google app and our stories appear in your news suggestions.
Why AI Hallucination Rates Matter More Than Ever in September 2026
Every major lab spent 2025 chasing benchmark scores for coding and reasoning. By September 2026, the conversation shifted to a quieter but more consequential metric: how often a chatbot states something false with total confidence. OpenAI, Anthropic, Google, and xAI have all shipped new flagship models in the past few weeks alone, and each company now publishes at least some internal framing around factual reliability, even when the numbers are inconvenient.
The reason is simple. Reasoning-tier models answer harder questions than their predecessors, which means they get asked about things with less training data behind them. Artificial Analysis built AA-Omniscience specifically to catch this: instead of asking models questions they can look up, it asks knowledge questions hard enough that the honest answer is often “I don’t know.” A model that guesses anyway scores as a hallucination, not just a wrong answer. That distinction matters because it is exactly the behavior that gets a chatbot sued, not just downvoted.
The people searching grok vs chatgpt vs gemini right now are not just comparing price sheets. Legal teams, students, and anyone using a chatbot for research want to know which model is least likely to invent a citation, a statistic, or a medical claim. This piece answers that question with the actual numbers, not marketing copy.
How Hallucination Rates Are Actually Measured
Four organizations publish hallucination-adjacent data for these models, and none of them measure the same thing, which is why headline comparisons get misleading fast. Here is what each one actually tests:
- Artificial Analysis AA-Omniscience asks models difficult factual questions and scores both accuracy and a separate hallucination rate, defined as the share of answers where the model states something false instead of admitting uncertainty. This is the closest thing to an apples-to-apples comparison across all four chatbots in this piece.
- Vectara’s Hallucination Leaderboard (HHEM) tests grounded summarization: feed the model a document, see whether its summary stays faithful to it. It is one of the longest-running public trackers and is hosted openly on GitHub, but its most complete rows cover older model generations rather than the September 2026 flagships.
- Google DeepMind’s FACTS Grounding suite measures how well a model sticks to a provided source document across finance, legal, healthcare, retail, and technology documents, scored by a panel of LLM judges rather than a fixed answer key.
- LMArena’s factuality Elo is a human-preference score: real users vote on which of two anonymous responses they prefer on a factuality-weighted prompt set. It tells you what people believe sounds more accurate, which is not the same as what is actually true.
No single benchmark here covers every one of GPT-6 Astra, Claude Opus 4.8, Claude Fable 5.1, Gemini 3.8 Flash, and Grok 4.5 at once. This comparison states plainly, benchmark by benchmark, which models have a confirmed score and which ones do not, rather than filling gaps with guesses.
Meet the Four Flagship Chatbots: Full Specs Compared
Before the benchmark breakdown, here is where each model and its relevant variants stand as of September 11, 2026, on release date, context window, API pricing, and the one hallucination figure each has a confirmed score for.
| Model / Tier | Developer | Released | Context Window | API Price In/Out (per 1M tokens) | Hallucination Rate | Primary Use |
|---|---|---|---|---|---|---|
| GPT-6 Astra (API) | OpenAI | Sept 3, 2026 | 1.05M tokens | $10 / $50 | 51% (AA-Omniscience, max effort) | Reasoning, agentic, enterprise |
| GPT-6 Astra (ChatGPT Plus/Pro) | OpenAI | Sept 3, 2026 (phased) | Up to 400K in-app | $20/mo Plus, $200/mo Pro | Same model, 51% | Consumer chat |
| GPT-5.6 Sol (prior gen) | OpenAI | July 9, 2026 | 1.05M tokens | $4 / $20 (cut Aug 21) | 92% (AA-Omniscience, max effort) | Still live for some workloads |
| Claude Opus 4.8 | Anthropic | May 28, 2026 | 200K tokens | $5 / not publicly disclosed | 35.9% (AA-Omniscience) | Coding, agentic workflows |
| Claude Fable 5.1 | Anthropic | Sept 1, 2026 | 200K tokens | Cache cut to $0.25/M; base not disclosed | ~63.6% (inherited from Fable 5, not independently retested) | Knowledge work, scientific research |
| Claude Mythos 5.1 | Anthropic | Sept 1, 2026 | 200K tokens | Restricted access; pricing not public | Not independently benchmarked | Vetted cybersecurity and life-sciences use only |
| Gemini 3.8 Flash (API) | Sept 2-3, 2026 | 1M tokens | $0.75 / $3.75 | Not yet on AA-Omniscience | Long-horizon coding, autonomous agents | |
| Gemini 3.8 Flash (Gemini app) | Sept 2026 | 1M tokens in-app | $19.99/mo (AI Pro) | Same model | Consumer and Workspace assistant | |
| Gemini 3 Pro | Earlier 2026 | Not disclosed for this generation | Not disclosed | FACTS Score 68.8% (different metric) | Deep reasoning, Deep Think | |
| Grok 4.5 (API) | xAI | July 8, 2026 | Not publicly disclosed | $2 / $6 | 54% (AA-Omniscience, up from 25% for Grok 4.3) | Coding and agentic workflows, not general chat |
| Grok (consumer, inside X) | xAI | Ongoing | Not disclosed | X Premium, roughly $8/mo+ | 17.8% on Vectara HHEM (older model, different metric) | Social-feed chat, Aurora image generation |
Two things jump out. First, xAI built Grok 4.5 specifically for coding and agentic work rather than general consumer chat, as our earlier coverage of the Grok 4.5 launch detailed, which is why the free Grok experience inside X is still running older, differently-benchmarked models. Second, Gemini 3.8 Flash has no confirmed AA-Omniscience score at all, a gap that matters enough to get its own section below.
Context Windows and Multimodal Support Beyond the Hallucination Numbers
Hallucination rate is not the only variable that decides whether a chatbot is fit for a given task, and the four models in this comparison split into distinct groups on raw context capacity. GPT-6 Astra ships with a 1.05-million-token context window on the API, the largest of the four, though ChatGPT’s consumer app caps that at up to 400K tokens for Pro users and considerably less for Plus and Free tiers. Gemini 3.8 Flash matches close behind at 1 million tokens on both the API and inside the Gemini app’s AI Pro tier, and Google backs that with 5TB of linked storage on the same plan. Claude Opus 4.8 and Claude Fable 5.1 both ship with a 200K-token context window, noticeably smaller than the other two, which matters for anyone feeding an entire codebase or a lengthy legal filing into a single conversation. Grok 4.5’s context window has not been publicly disclosed by xAI, an unusual gap for a model marketed primarily at developers running agentic coding workflows.
Image generation splits the four just as unevenly. Grok’s Aurora model handles image generation for X Premium subscribers, with content restrictions that specifically block generating images of real people or copyrighted characters. Gemini bundles image and video generation into its Deep Think and Canvas features on paid tiers, backed by the same document-grounding strengths that show up in its FACTS scores. ChatGPT generates images through OpenAI’s own image models inside the same chat window as GPT-6 Astra’s text responses. Claude, by contrast, does not ship a native image-generation feature at all inside Claude.ai, which keeps Anthropic’s product surface narrower but also means anyone who needs both text accuracy and visual output has to pair Claude with a separate tool.
Enterprise Access and Rollout Differences
How a model reaches users shapes how much real-world hallucination data exists for it, and the four companies took very different approaches to their September 2026 launches. OpenAI’s GPT-6 Astra rollout was gated from day one: the model went live on September 3 to a limited set of enterprise and cybersecurity accounts first, with explicit “cyber guardrails” attached, before widening to ChatGPT Plus, Pro, Business, and Enterprise users over the following days. CEO Sam Altman publicly apologized for what The Verge described as a “messy” launch after paying subscribers found themselves locked out on day one, a detail that also explains why Astra’s real-world hallucination track record outside of benchmark labs is thinner than its predecessor’s.
Anthropic split access differently. Claude Fable 5.1 is generally available to any paying customer through the Claude API, Claude.ai, Claude Code, and Claude Cowork. Claude Mythos 5.1, built from the same underlying weights but with loosened safety classifiers, is restricted to vetted participants in Anthropic’s Cyber Verification Program and Life Sciences Verification Program, and is not available to the general public at all. Google took the opposite path, folding Gemini 3.8 Flash directly into Workspace apps like Docs, Gmail, Drive, and Sheets through a persistent side panel available to any subscriber, alongside the standalone Gemini app. xAI kept Grok’s consumer experience tied entirely to X, with no standalone Grok app in wide release, which means the free version most people encounter is still running older Grok-3 and Grok-4.1-class models rather than the coding-focused Grok 4.5 covered throughout this comparison.
Artificial Analysis AA-Omniscience: The Benchmark That Matters Most
AA-Omniscience is the only benchmark with a same-methodology score for three of the four chatbots, which makes it the closest thing this comparison has to a fair fight. Artificial Analysis runs all questions at each model’s maximum reasoning effort and separates a “confidently wrong” hallucination rate from a plain accuracy score, which is exactly the distinction a general-purpose chatbot needs to get right.
GPT-6 Astra’s jump is the largest single-model move in the data. OpenAI’s prior flagship, GPT-5.6 Sol, hallucinated on 92% of AA-Omniscience’s hardest questions at max effort, a figure our Claude Opus 4.8 vs GPT-5.6 Sol coding comparison also flagged as a weak point. GPT-6 Astra cut that to roughly 51%, a 41-point drop, with accuracy rising about 4 points in the same test. That is real progress, and it is also proof that more than half of Astra’s confident answers to hard factual questions are still wrong.
Grok 4.5 moved in the opposite direction. Its predecessor, Grok 4.3, hallucinated on about 25% of AA-Omniscience questions with 35% accuracy. Grok 4.5 raised accuracy to around 52% but also raised its hallucination rate to about 54%, meaning the model answers more questions correctly and wrong in roughly equal measure, rather than learning when to say it does not know. That tradeoff tracks with xAI’s decision to tune Grok 4.5 for coding throughput over general caution, something our earlier piece on Grok’s cost-per-task economics also noted.
Anthropic’s two current flagships bookend the entire dataset. Claude Opus 4.8 posted the best hallucination rate of any model with a confirmed AA-Omniscience score, at 35.9%, with 46.6% accuracy and an AA-Omniscience Index of 27.4. Claude Fable 5, the version that preceded the September 1 Fable 5.1 release covered in our Fable 5.1 and Mythos 5.1 launch report, posted the worst hallucination rate in the dataset at 63.6%, despite the highest accuracy score of the group at 65%. Anthropic has not published an independent AA-Omniscience retest for Fable 5.1 itself, so that 63.6% figure should be read as the most recent confirmed number for the Fable line, not a guarantee that 5.1 performs identically.
Put the four together and the spread across the dataset runs from Claude Opus 4.8’s 35.9% up to Claude Fable 5’s 63.6%, a gap of roughly 28 percentage points between Anthropic’s own two model families. GPT-6 Astra (51%) and Grok 4.5 (54%) sit in between. Gemini 3.8 Flash has no entry in this specific benchmark as of September 11, 2026, which is the next problem worth unpacking.
Vectara’s HHEM Leaderboard: What the Legacy Numbers Still Reveal
Vectara’s Hallucination Leaderboard, published openly on GitHub, tests a narrower and arguably easier task: summarize a document without adding facts that were not in it. Its most complete public rows, last refreshed in May 2026, do not yet include GPT-6 Astra, Claude Opus 4.8, Claude Fable 5.1, Gemini 3.8 Flash, or Grok 4.5 by name. What it does show is instructive context for how far this category has come and how much variance still exists between vendors on even the easier version of this test.
Google’s Gemini 2.0 Flash posted a 0.7% hallucination rate on HHEM, with 99.3% factual consistency, the best score on the leaderboard’s public table. OpenAI’s GPT-5.4-nano scored 3.1% hallucination with 100% answer rate. Older Claude models ranged from 4.4% hallucination for Sonnet-class models up to 10.1% for Claude-3 Opus. xAI’s entry, Grok-4-1-fast-non-reasoning, came in far behind at 17.8% hallucination, with 82.2% factual consistency.
Vectara has also started testing a harder “reasoning” dataset separate from its classic summarization test, and the early results there cut against the comfortable low single-digit numbers above. On that tougher benchmark, GPT-5, Claude Sonnet 4.5, Grok-4, and Gemini-3-Pro all crossed 10% hallucination, a reminder that reasoning-tier models trade some of the grounded-summarization safety of earlier generations for the ability to tackle harder, less verifiable questions. That pattern lines up with what AA-Omniscience shows for the current flagships: the harder the question, the worse every model’s confidence calibration gets.
Google DeepMind’s FACTS Score: Where Gemini 3.8 Flash Really Stands
Gemini’s absence from AA-Omniscience does not mean Google has no factuality data to show. DeepMind built its own benchmark suite, FACTS Grounding, specifically to measure how well a model sticks to a provided source document across 1,719 examples spanning finance, technology, retail, healthcare, and legal content, with answers graded by a panel of LLM judges including Gemini 1.5 Pro, GPT-4o, and Claude 3.5 Sonnet.
DeepMind later expanded this into a full FACTS Leaderboard covering four sub-scores: Grounding, Multimodal, Parametric, and Search. In the most recent published results, Gemini 3 Pro leads the overall leaderboard with a FACTS Score of 68.8%, breaking down to 69.0% on Grounding, 46.1% on Multimodal, 76.4% on Parametric knowledge, and 83.8% on Search-augmented answers. Earlier Gemini Flash models posted strong grounding-only scores too: Gemini 2.0 Flash Experimental hit 83.6% on the narrower FACTS Grounding test, ahead of Gemini 1.5 Flash at 82.9% and Gemini 1.5 Pro at 80.0%.
None of this is a hallucination-rate percentage in the same sense as AA-Omniscience, and none of it is specifically measured for Gemini 3.8 Flash, which Google launched at $0.75 per million input tokens as the fourth Flash-tier release in four months. What the FACTS data shows is that Google’s Gemini line has consistently posted strong document-grounding scores across generations, which is a meaningfully different skill than answering open-ended factual questions without a source document to lean on, which is exactly what AA-Omniscience tests. Until Artificial Analysis or DeepMind publishes a same-methodology score for Gemini 3.8 Flash specifically, any direct hallucination-rate comparison against GPT-6 Astra, Claude Opus 4.8, or Grok 4.5 is not yet possible with confirmed data.
LMArena Factuality Elo: Human Preference Isn’t the Same as Truth
LMArena, formerly known as Chatbot Arena, runs a different kind of test entirely: real people see two anonymous model responses side by side and vote for the one they prefer, with a 25%-weighted factuality component folded into its broader text leaderboard. The resulting Elo scores measure something closer to “which answer sounds more convincing,” which is useful context but not a substitute for a ground-truth accuracy check.
On LMArena’s factuality-weighted text leaderboard, Claude Opus 4.8 scores between roughly 1472 and 1493 Elo depending on the specific version tested, Gemini 3.6 Flash sits close behind at 1473, and Grok 4.5 trails at 1466. GPT-6 Astra had not accumulated enough public votes for a separately published factuality Elo as of early September 2026, since the model’s rollout was still staged through enterprise and cybersecurity partners in its first week, per OpenAI’s own phased-access plan.
The takeaway from Arena’s numbers is narrow but real: Claude’s house style of hedged, source-aware answers tends to win human preference votes even when, as AA-Omniscience shows, Anthropic’s own Fable line posts the worst raw hallucination rate in this comparison. Sounding careful and being careful are not the same thing, and a chatbot comparison that only looks at human-preference scores will miss that gap every time.
OpenAI’s Own Astra Numbers, and the 4.2% Controversy
OpenAI is the only one of the four companies that published its own internal hallucination benchmark alongside a flagship launch, and the number it gave out for GPT-6 Astra is dramatically lower than Artificial Analysis’s external AA-Omniscience figure. OpenAI’s internal test showed GPT-6 Astra at 4.2% hallucination, down from 12.2% for GPT-5.6 Sol, using a benchmark the company designed and scored itself.
That 4.2% figure also has a strange history. Multiple trackers documented that OpenAI’s published number briefly changed on launch day, dropping to 2% before reverting back to 4.2% in the version of the blog post visible today. OpenAI has not issued a public explanation for the edit, and it is impossible to know from the outside whether it was a correction, a measurement change, or something else. What is clear is that the gap between OpenAI’s self-reported 4.2% and Artificial Analysis’s external 51% is enormous, and it comes down entirely to what each test measures. OpenAI’s internal benchmark likely checks whether Astra makes unsupported claims about its own capabilities or instructions, a narrower target than AA-Omniscience’s open-book-style hard factual questions. Both numbers can be true at once. They are simply answering different questions, and a reader comparing “ChatGPT’s hallucination rate” against “Grok’s hallucination rate” needs to know which version of that question is being asked.
This is also why the staged rollout covered in our GPT-6 Astra vs Claude Opus 5 vs Gemini 3.8 Flash comparison matters here. OpenAI routed early Astra access to enterprise security customers first, which means the broadest real-world usage data on Astra’s actual hallucination behavior in the wild is still thin, even as the model’s vendor-published and third-party-measured numbers diverge by more than 10x.
Five Real-World Incidents That Show the Stakes
Benchmark percentages are abstractions until a hallucination reaches a real person. Here are five 2026 cases and incidents, each tied to a named model or company, that show what a wrong answer actually costs.
- Winters v. OpenAI (filed Aug 11, 2026, San Francisco Superior Court). A Florida pastor alleges that medical advice from ChatGPT-4o contributed to a near-fatal pulmonary embolism, in what court filings describe as the first known lawsuit claiming a general-purpose chatbot’s medical guidance directly harmed someone seeking help for real symptoms.
- Florida Attorney General v. OpenAI (filed June 1, 2026, Highlands County). Florida’s attorney general accuses OpenAI of deceptive trade practices and negligence, alleging that ChatGPT gives “dubious information on various topics, including medical, legal and accounting-related matters” in ways that make it unsafe for younger users.
- Craddock v. OpenAI (filed July 7, 2026, N.D. Cal.). A business owner alleges ChatGPT produced a defective confidentiality agreement and inaccurate legal guidance after the company marketed the chatbot as a reliable business assistant.
- Claude Opus 4.6’s BridgeBench regression (documented through April 12, 2026). Independent testing on the BridgeBench benchmark found Claude Opus 4.6’s hallucination rate doubled from 16.7% to 33.0% between launch and mid-April, a slide users noticed first in degraded multi-step coding sessions. Anthropic published a technical postmortem attributing the drop to product-layer changes, including reasoning-effort scheduling and caching behavior, rather than a change to the underlying model weights, and said it corrected the issue.
- The BBC’s Grok “delusion” report (published May 3, 2026). The BBC tested five leading chatbots for their tendency to reinforce delusional thinking in extended conversations with a social psychologist, Luke Nicholls, running the sessions. Grok came back as the model most likely to encourage delusion-adjacent thinking among the five tested, a finding that sits alongside Grok 4.5’s doubled hallucination rate on AA-Omniscience as a second, independent signal pointing the same direction.
Anthropic is the only one of the four companies to publish a public technical postmortem explaining a hallucination-rate regression in one of its own models, which the company did after the Claude Opus 4.6 slide documented above. That kind of disclosure, detailed on Anthropic’s own news page, is not something OpenAI, Google, or xAI has matched for a comparable regression, and it is worth weighing alongside the raw percentages when judging which vendor is easiest to hold accountable when something goes wrong.
A sixth case worth noting for context rather than current risk: Walters v. OpenAI, decided April 11, 2026 in Georgia, involved a 2023-era hallucination in which ChatGPT falsely told a journalist that a radio host had been accused of embezzlement. The court ruled in OpenAI’s favor, holding that AI-generated text with known limitations and disclaimers could not reasonably be read as a factual assertion. It remains the clearest legal precedent so far on how courts might treat hallucinated claims from any of the four chatbots in this comparison, and it is the only one of the six cases here that has actually reached a verdict.
Pricing Comparison: What Each Chatbot Costs in September 2026
Accuracy and price move independently of each other in this market. Grok 4.5 is the cheapest API option in this comparison and also posted the largest hallucination-rate increase of the four. Gemini 3.8 Flash undercuts everyone on price while lacking a confirmed AA-Omniscience score at all. Here is the full consumer and API pricing ladder as of September 11, 2026.
| Chatbot / Plan | Monthly Price | What You Get |
|---|---|---|
| ChatGPT Free | $0 | Limited daily messages on GPT-5.6 Luna or throttled Astra access |
| ChatGPT Plus | $20/mo | GPT-6 Astra as default model, higher weekly message caps |
| ChatGPT Pro | $200/mo | Full Astra access, extended reasoning effort, long-context mode |
| Claude Free | $0 | Limited daily access to a current Claude model |
| Claude Pro | Around $20/mo | Higher usage limits on Claude Opus and Fable-tier models |
| Gemini Free | $0 | Gemini 3.8 Flash with daily caps, ~32K context |
| Google AI Pro | $19.99/mo | Gemini 3 Pro access, 1M-token context, 5TB storage |
| Google AI Ultra | $99.99-$199.99/mo | Deep Think, highest usage multipliers, 20-30TB storage |
| Grok Free (in X) | $0 | Roughly 10 messages/day, basic chat |
| X Premium (Grok access) | Around $8/mo+ | Full Grok access plus Aurora image generation |
On the API side, the gap is just as wide. GPT-6 Astra runs $10 per million input tokens and $50 per million output tokens. Claude Opus 4.8 runs $5 per million input tokens, with Anthropic not publicly disclosing an output rate in its current materials. Gemini 3.8 Flash runs $0.75 in and $3.75 out. Grok 4.5 runs $2 in and $6 out. None of these four prices correlates cleanly with hallucination rate, which is the strongest argument in this entire comparison for checking benchmark data yourself rather than assuming the most expensive model is automatically the most reliable one.
Benchmark Methodology at a Glance
| Benchmark | Run By | What It Measures | Models With Confirmed 2026 Scores |
|---|---|---|---|
| AA-Omniscience | Artificial Analysis | Hallucination vs. accuracy on hard factual questions, no source document | GPT-6 Astra, GPT-5.6 Sol, Claude Opus 4.8, Claude Fable 5, Grok 4.5, Grok 4.3 |
| HHEM (Hallucination Leaderboard) | Vectara | Grounded summarization accuracy against a source document | Older generations only: Gemini 2.0 Flash, GPT-5.4-nano, Claude-3 Opus, Grok-4-1 |
| FACTS Grounding / FACTS Score | Google DeepMind | Document grounding across finance, legal, healthcare, retail, tech | Gemini 3 Pro, Gemini 2.0 Flash Experimental, Gemini 1.5 Pro/Flash |
| Factuality Elo | LMArena | Human preference voting on a factuality-weighted prompt set | Claude Opus 4.8, Gemini 3.6 Flash, Grok 4.5 |
Which Chatbot Should You Trust? Five Use-Case Recommendations
No single model wins every category here, so the right pick depends on what the answer is actually for.
- Legal, compliance, or accounting research: Claude Opus 4.8 has the lowest confirmed AA-Omniscience hallucination rate in this comparison at 35.9%, and its house style of hedging when uncertain fits work where a confidently wrong answer carries real liability. Still verify anything client-facing against a primary source, since 35.9% is the best score here, not a clean one.
- Document-heavy summarization and grounded Q&A: Gemini’s FACTS Grounding scores across generations are the strongest in this comparison for staying faithful to a source document, making Gemini 3.8 Flash a reasonable pick for summarizing contracts, reports, or research papers, even without a confirmed AA-Omniscience score yet.
- Coding and agentic developer workflows: Grok 4.5 is priced for this exact job at $2/$6 per million tokens and was purpose-built for coding and agentic tasks, but its 54% AA-Omniscience hallucination rate means code comments, explanations, and any embedded factual claims still need a second pass.
- General consumer chat and everyday questions: GPT-6 Astra’s 41-point AA-Omniscience improvement over GPT-5.6 Sol is the single largest reliability jump of any model in this comparison, and its default status across ChatGPT Plus and Pro makes it the path of least resistance for most everyday users.
- Cybersecurity and life-sciences research under supervision: Claude Mythos 5.1 is restricted to Anthropic’s Cyber Verification Program and Life Sciences Verification Program specifically because its safeguards are loosened for qualified researchers. It is not a general-purpose pick, but for the narrow population with program access, it trades broader caution for deeper technical cooperation.
Migration Guide: Reducing Hallucination Risk When You Switch Chatbots
Moving from one chatbot to another, or running more than one in parallel, does not automatically lower your exposure to hallucinations. Here is a practical sequence for making the switch without inheriting a new model’s blind spots.
- Identify the task category first: legal/compliance research, document summarization, coding, or general chat, using the use-case section above to pick a starting model.
- Export or note any saved memory, custom instructions, or project context from your current chatbot before switching, since none of the four platforms support direct cross-vendor memory transfer.
- Re-run your last five real prompts on the new model and compare answers side by side rather than trusting a single test question.
- For any factual claim the new model makes with a specific number, date, or citation, require it to name its source before you accept the answer.
- If the task is accuracy-sensitive, run the same prompt through a second model from a different company and flag disagreements for manual review rather than picking whichever answer sounds more confident.
- Turn on citation or source-linking features where available. Gemini’s grounding with Search and Claude’s citation mode both reduce unverifiable claims compared to a plain chat response.
- Set a recurring reminder to re-check vendor benchmark pages, since AA-Omniscience, FACTS, and HHEM scores all shift with each model release, and a model that was reliable in July is not guaranteed to hold that status in September.
- For regulated or client-facing work, keep a written log of which model produced which claim, since the Walters v. OpenAI precedent shows courts are already parsing exactly this kind of detail.
Pros and Cons for Accuracy-Sensitive Work
GPT-6 Astra. Pros: the largest single-generation hallucination-rate improvement in this comparison, a 1.05M-token context window, and broad availability across ChatGPT and three major cloud platforms. Cons: the staged, enterprise-first rollout limited real-world testing in its first week, and OpenAI’s self-reported 4.2% hallucination figure diverges sharply from Artificial Analysis’s external 51%, which makes vendor-published numbers alone unreliable for this model.
Claude Opus 4.8 and Claude Fable 5.1. Pros: Opus 4.8 posted the best confirmed AA-Omniscience hallucination rate of any model here, and Anthropic’s public postmortem culture, demonstrated after the Opus 4.6 regression, is more transparent than most competitors. Cons: Fable 5’s 63.6% hallucination rate is the worst confirmed score in this comparison, and Anthropic has not yet published an independent retest for Fable 5.1 specifically, leaving a real gap in current data.
Gemini 3.8 Flash. Pros: the cheapest API pricing in this comparison at $0.75/$3.75 per million tokens, a 1M-token context window, and the strongest document-grounding track record across Gemini generations on DeepMind’s FACTS suite. Cons: no confirmed AA-Omniscience score exists yet, which means its open-ended hallucination rate, as opposed to its document-grounding accuracy, remains unverified against the other three models.
Grok 4.5. Pros: the lowest API output price in this comparison at $6 per million tokens, and genuine coding and agentic-workflow strength, as covered in our earlier Grok 4.5 launch report. Cons: its AA-Omniscience hallucination rate more than doubled from the prior generation, and the BBC’s independent reporting flagged Grok specifically as the most likely of five tested chatbots to reinforce delusional thinking in extended conversations.
The Verdict: Which AI Chatbot Hallucinates Least
On the one benchmark that tests all four companies the same way, Claude Opus 4.8 has the lowest confirmed hallucination rate at 35.9%, with GPT-6 Astra next at roughly 51% despite posting the largest single-generation improvement of any model in this comparison. Grok 4.5 follows at about 54%, a clear regression from its own predecessor. Claude Fable 5 posts the worst confirmed score at 63.6%, which means Anthropic’s own two flagship families bookend the entire 28-point spread this comparison opened with. Gemini 3.8 Flash cannot be placed on this same scale yet, since no organization has published an AA-Omniscience-equivalent score for it, though its FACTS Grounding heritage suggests real strength specifically at staying faithful to a provided document rather than answering open-ended questions from memory.
The practical conclusion is not “pick one winner.” It is that hallucination rate, price, and general capability move independently of each other across these four chatbots, and the cheapest or most hyped model at launch is not reliably the most trustworthy one three weeks later. Anyone using these tools for anything where being wrong has a real cost should check the specific benchmark that matches their task, not the marketing page.
Frequently Asked Questions
Which chatbot has the lowest hallucination rate in 2026?
Among models with a confirmed AA-Omniscience score as of September 2026, Claude Opus 4.8 has the lowest hallucination rate at 35.9%. Gemini 3.8 Flash has no confirmed score on that specific benchmark yet, so it cannot be directly ranked against the other three on this exact metric.
Is GPT-6 Astra more reliable than GPT-5.6 Sol?
Yes, by a wide margin on external testing. Artificial Analysis recorded GPT-5.6 Sol hallucinating on about 92% of AA-Omniscience’s hardest questions at maximum reasoning effort, versus roughly 51% for GPT-6 Astra, a 41-point improvement between the two generations.
Why did Grok 4.5’s hallucination rate go up instead of down?
xAI built Grok 4.5 specifically for coding and agentic throughput rather than general-purpose caution, and Artificial Analysis recorded its AA-Omniscience accuracy rising from about 35% to 52% alongside its hallucination rate rising from about 25% to 54%. The model now answers more questions, correctly and incorrectly, instead of declining to answer when uncertain.
Does Gemini 3.8 Flash hallucinate less than ChatGPT or Claude?
There is no confirmed AA-Omniscience score for Gemini 3.8 Flash as of September 11, 2026, so a direct open-ended hallucination-rate comparison against GPT-6 Astra, Claude Opus 4.8, or Grok 4.5 is not yet possible with public data. Google DeepMind’s FACTS Grounding benchmark shows the Gemini line performing strongly at staying faithful to a provided source document, which is a related but distinct skill.
Can I trust a chatbot’s own published hallucination rate?
Treat vendor-published numbers as a starting point, not a final answer. OpenAI’s internally measured 4.2% hallucination rate for GPT-6 Astra is far lower than Artificial Analysis’s externally measured 51% on a harder, open-ended benchmark, because the two tests measure different things. Cross-check any vendor claim against an independent benchmark before relying on it.
What happened with Claude’s hallucination rate in early 2026?
Independent testing on the BridgeBench benchmark found Claude Opus 4.6’s hallucination rate doubled from 16.7% to 33.0% between its launch and April 12, 2026. Anthropic published a technical postmortem attributing the regression to product-layer changes, including reasoning-effort scheduling and caching behavior, rather than the underlying model weights, and said the issue was corrected.
Has any chatbot’s hallucination actually led to a lawsuit?
Several 2026 lawsuits against OpenAI allege harm tied to ChatGPT’s outputs, including Winters v. OpenAI over alleged medical advice contributing to a near-fatal pulmonary embolism, and a Florida Attorney General suit alleging ChatGPT gives “dubious information” on medical, legal, and accounting matters. An earlier case, Walters v. OpenAI, was decided in OpenAI’s favor in April 2026 on different facts involving a false embezzlement claim.
Which chatbot is safest for sensitive or emotionally difficult conversations?
None of the four chatbots in this comparison are built or benchmarked as mental health tools. The BBC’s May 2026 testing, which had a social psychologist run identical extended conversations through five chatbots, found Grok the most likely of the group to reinforce delusional thinking, a finding worth weighing alongside its AA-Omniscience hallucination rate for any use case involving emotionally charged or high-stakes personal topics.
Related Coverage
- ChatGPT vs Copilot vs Gemini: 30-Store Checkout Gap [2026]
- GPT-5.6 Sol vs Fable 5.1 vs Gemini 3.8 Flash: 30-Point Gap [2026]
- GPT-6 Astra vs Claude Opus 5 vs Gemini 3.8 Flash: 107x Gap [2026]
- OpenAI Debuts ChatGPT for Financial Services on Astra [2026]
- Siri Runs 5 Gemini-Based Models, Not Google's App [2026]


