Fact-Checking Using AI

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Large language models (LLMs) can both correct inaccurate beliefs (1, 2, 3) and manipulate attitudes (4, 5, 6). As LLMs are increasingly deployed as fact-checkers on social media, their impact depends not only on the models’ technical capability, but also on the choices of their designers and the attitudes and behaviors of platform users. Here, we shed light on all three dimensions of this sociotechnical system. We analyze an exhaustive sample of 1,555,281 English-language fact-check requests sent to Grok and Perplexity bots on X from February through September 2025 and a random sample of 184,124 fact-checks from Grok from September 2025 to August 2026. We also benchmark 19 frontier LLMs against professional fact-checkers’ evaluations of hundreds of news headlines, and conduct a preregistered survey experiment (N=1,592) to evaluate correction efficacy. On the user side, Grok requesters skew heavily Republican; familiarity with and trust in specific LLMs vary substantially, such that trust in AI is no longer a unitary construct; and revealing that a fact-check comes from Grok reduces belief updating among Democrats while increasing belief updating among Republicans. Examining LLM outputs, we find that from a technical perspective, LLM fact-checks can match or exceed agreement among professional fact-checkers. Yet we also show that this capacity is readily undermined by designers’ choices: for example, restricting sources (e.g. Truth Social AI) or disabling search can degrade accuracy. Designers and users, more so than technical capability, are now the primary binding constraints.

Article activity feed