
Are AI chatbots reliable against disinformation? A study published on August 30, 2026 by NPR and NewsGuard gives a numeric, nuanced answer. Across 30 questions built from false narratives spread by Russia, China and Iran, six mainstream chatbots (ChatGPT, Gemini, Copilot, Meta AI, Grok and Claude) correctly debunked those narratives about 75% of the time. For an SMB leader who increasingly relies on these tools for market or regulatory watch, the question is anything but theoretical.
At a glance
- On August 30, 2026, NPR and NewsGuard published the results of a test run on six AI chatbots (ChatGPT, Gemini, Copilot, Meta AI, Grok, Claude) against 15 false narratives spread by Russia, China and Iran between December 2025 and July 2026 (source: NPR).
- On average, the chatbots correctly debunked these narratives about 75% of the time (source: NPR/NewsGuard).
- AI-generated search summaries (Google, Bing, DuckDuckGo) performed worse than the chatbots: Google's remains the most reliable, Bing's fails most often.
- Another NewsGuard indicator, the AI False Claims Monitor, points to a more worrying long-term trend: on broader, controversial news topics, the average rate of false claims reportedly rose from 18% to 35% between fall 2024 and fall 2025 (source: Axios).
- For an SMB, the takeaway is twofold: mainstream chatbots are generally solid on well-documented geopolitical topics, but vigilance is still needed on breaking news and controversial subjects.
A new study: how NPR and NewsGuard tested AI chatbots
NewsGuard is an organization that rates the reliability of news sources and regularly audits major AI models. For this investigation, its researchers first identified 15 false narratives spread by Russian, Chinese and Iranian state actors between December 2025 and July 2026. For each one, they wrote two questions: a neutral question ("did this happen?") and a leading question ("why did this happen?"), for 30 questions in total.
These 30 questions were then put to six chatbots (ChatGPT, Gemini, Copilot, Meta AI, Grok, Claude) as well as four search engines (Google, Bing, DuckDuckGo, Yandex), analyzing the AI-generated summaries shown at the top of results for the first three. Data was collected in mid-July 2026, then the answers were compared against fact-checks produced by NewsGuard's researchers.
Definition
NewsGuard's AI False Claims Monitor is a monthly barometer, launched in July 2024, that tests eleven chatbots on broadly controversial news topics. It does not measure the same thing as the NPR/NewsGuard study from August 2026: the latter focuses on precisely documented state-sponsored false narratives, an easier case for an AI to resolve.
The results: three false claims out of four caught
The headline result is encouraging: on average, the six chatbots correctly caught the false narratives about 75% of the time. As digital literacy expert Mike Caulfield told NPR: "if an educator gave their students a similar assignment using a traditional search engine and saw three-quarters of them getting the answers right, you would be ecstatic."
AI-generated search summaries, on the other hand, did worse.
| Tool tested | Observed performance | Source |
|---|---|---|
| AI chatbots (average of 6 tested) | About 75% correct debunking rate | NPR/NewsGuard, August 2026 |
| Google's AI summary (AI Overview) | The most reliable summary; shown for 27 of 30 queries, but about 1 claim in 9 unsupported by its cited source | NPR/NewsGuard, August 2026 |
| DuckDuckGo's AI summary | Intermediate results, between Google and Bing | NPR/NewsGuard, August 2026 |
| Bing's AI summary | The least reliable; shown for under half of queries, and a majority failure to debunk when it did appear | NPR/NewsGuard, August 2026 |
Chatbots vs. search engine AI summaries: who wins?
AI chatbots (ChatGPT, Gemini, Claude...)
Queried directly, with more extended reasoning on the question asked. Highest correct debunking rate in the study, about 75% on average. Still has room to improve on leading questions.
Search engine AI summaries
Generated at the top of search results, often from a limited set of indexed sources. More uneven performance: Google leads, DuckDuckGo is in between, Bing lags. A summary can also fail to appear for a chunk of queries.
This gap is partly explained by how the tools are designed. A conversational chatbot can draw on more extended reasoning and weigh several angles before answering. A search engine summary, by contrast, has to produce a quick synthesis from the top indexed results, making it more dependent on the quality (and sometimes the reliability) of whatever ranks highest.
The flip side: a broader trend worth watching
A necessary caveat
The 30 questions in the NPR/NewsGuard study cover specific geopolitical events that are otherwise well documented. Google, quoted by NPR, points out that such cases remain "rare" compared to everyday chatbot use. The strong score here does not guarantee the same reliability on every topic, especially ones still being debated or thinly covered by reliable sources.
That is exactly what another NewsGuard indicator suggests: the AI False Claims Monitor, which has tracked eleven chatbots monthly on broadly controversial news topics since July 2024. According to figures reported by Axios, the average rate of false claims found rose from 18% to 35% between fall 2024 and fall 2025, partly because more chatbots now have real-time web search, which also exposes them to lower-quality sources circulating in the moment. The two studies don't measure exactly the same thing, but reading them together is useful: an AI chatbot can be solid on a well-documented geopolitical narrative and still be vulnerable on fast-moving, unsettled news.
What this means for your SMB
Cross-check with at least one other source
Before basing a business decision or external communication on an AI answer, verify it against a second source, ideally a recognized outlet or institution.
Prefer a neutral question over a leading one
The study suggests a question that already presupposes an explanation ("why did X happen?") is riskier than a factual one ("did X happen?"). Phrase your queries neutrally.
Be more cautious with search engine summaries
On a sensitive topic, a conversational chatbot that cites sources is, according to this study, more reliable than an AI summary shown at the top of search results.
Train your teams to verify
An internal AI usage policy (see our article on shadow AI) should include a simple rule: any sensitive information relayed by an AI gets checked before it is shared.
This finding fits into a wider conversation about trust in generative AI at work. LUWAI has already covered how AI vendors are strengthening their safeguards, for instance with Claude's content watermark, or how the industry is organizing collectively against AI-driven cyberattacks. Information reliability is a neighboring, equally important challenge for any SMB using AI day to day.
FAQ
Can you trust ChatGPT or Gemini to verify information?
Broadly, yes for well-documented geopolitical facts: the NPR/NewsGuard study from August 2026 found an average correct debunking rate of about 75% across six chatbots tested. That confidence should stay measured on recent, controversial, or thinly covered topics, where NewsGuard's AI False Claims Monitor finds higher error rates.
What's the difference between an AI chatbot and a search engine's AI summary?
A conversational chatbot (ChatGPT, Gemini, Claude) answers a question by drawing on more extended reasoning. A search engine's AI summary (Google AI Overview, Bing) quickly synthesizes the top indexed results for a query. According to the NPR/NewsGuard study, chatbots are generally more reliable than search engine summaries on this specific exercise.
Have AI chatbots gotten more reliable since 2024?
Results vary by indicator. On specific, well-documented geopolitical narratives, the August 2026 study shows a solid score (75%). But on broadly controversial news topics, NewsGuard's AI False Claims Monitor points to a decline, from 18% to 35% false claims between fall 2024 and fall 2025 (source: Axios).
How should an SMB use AI for its market watch without risking disinformation?
Systematically cross-check an AI answer against a second reliable source before any decision or external communication, favor neutral questions over leading ones, and train your teams on this verification habit through your AI usage policy.
Conclusion
AI chatbots are neither infallible nor systematically misleading: the NPR/NewsGuard study from August 2026 shows real but imperfect reliability against state-sponsored disinformation, clearly ahead of search engine summaries. For an SMB, the right practice is the same as for any information source: verify before deciding or communicating. To go further on responsible AI use at work, check out our other LUWAI Mag resources.


