AI-generated content in Wikipedia - a tale of caution

Mathias Schindler

39th Chaos Communication Congress (39C3): Power Cycles · Day 1 · Saal Ground

Overview

In an era increasingly shaped by artificial intelligence, Mathias Schindler, a veteran Wikipedian and co-founder of Wikimedia Germany, delivered a crucial talk at 39C3 titled "AI-generated content in Wikipedia - a tale of caution." This presentation peeled back the layers of a growing problem: the silent infiltration of Large Language Model (LLM)-hallucinated content into one of the world's most trusted knowledge repositories. Schindler's work, initially focused on identifying human errors in ISBNs, inadvertently uncovered a more insidious form of misinformation, where LLMs generate plausible-sounding but entirely fabricated literature references.

Watch on YouTube

Visual summary for AI-generated content in Wikipedia - a tale of caution by Mathias Schindler
Visual summary for AI-generated content in Wikipedia - a tale of caution by Mathias Schindler

Key moments

  1. 0:00 Emotional introduction and thanks to the CCC community
  2. 2:40 Speaker's background, current AI work, and caution
  3. 3:30 Wikipedia's philosophy: start research, not end research
  4. 6:00 Developing an ISBN checksum checker for Wikipedia
  5. 9:00 Discovery of a puzzling third category of unfindable references
  6. 10:00 The 'Aha!' moment: literature references hallucinated by ChatGPT

AI-generated content in Wikipedia - a tale of caution

Speakers: Mathias Schindler

Conference: 39C3

YouTube: https://www.youtube.com/watch?v=fKU0V9hQMnY

Overview

In an era increasingly shaped by artificial intelligence, Mathias Schindler, a veteran Wikipedian and co-founder of Wikimedia Germany, delivered a crucial talk at 39C3 titled "AI-generated content in Wikipedia - a tale of caution." This presentation peeled back the layers of a growing problem: the silent infiltration of Large Language Model (LLM)-hallucinated content into one of the world's most trusted knowledge repositories. Schindler's work, initially focused on identifying human errors in ISBNs, inadvertently uncovered a more insidious form of misinformation, where LLMs generate plausible-sounding but entirely fabricated literature references.

The talk highlights a critical challenge to the integrity of online encyclopedic projects like Wikipedia. As LLMs are increasingly used, sometimes without explicit declaration or understanding of their limitations, they introduce "anti-knowledge" – content that appears factual but is fundamentally false. Schindler's findings resonate beyond Wikipedia, impacting libraries, research institutions, and any domain reliant on verified information. This discussion underscores the urgent need for robust detection mechanisms, clear community policies, and a collective re-evaluation of human responsibility in the age of generative AI.

This issue is particularly significant given that Wikipedia itself serves as a foundational training dataset for many LLMs. The introduction of synthetic, erroneous content back into Wikipedia risks "poisoning the well," creating a dangerous feedback loop where models are trained on increasingly unreliable data, further perpetuating misinformation. Schindler's address serves as a stark warning and a call to action for developers, users, and policymakers to confront the ethical and technical implications of AI-generated content on our shared knowledge infrastructure.

Background

▶ Watch: Emotional introduction and thanks to the CCC community (0:00)

Wikipedia has long operated under the guiding principle that it is "a good place to start research, [but] definitely a terrible place to end research." This statement implicitly places a moral obligation on Wikipedians to provide reliable resources for users to continue their research elsewhere, with literature references being a cornerstone of this commitment. These references are not merely footnotes; they are an integral part of Wikipedia's verifiable knowledge base, with the platform offering features like automatic ISBN conversion to facilitate access to external materials from national libraries, bookstores, and other sources.

Central to Schindler's investigation is the International Standard Book Number (ISBN). ISBNs possess inherent verification mechanisms, including a checksum as their last digit. This checksum is calculated using specific formulas (for ISBN-10 and ISBN-13) which ensure that common errors like single-digit mistakes or transpositions of two numbers will result in a mismatched checksum. This mathematical robustness makes ISBNs a seemingly reliable identifier for literature.

Schindler's journey began with a personal project: to build an ISBN checksum checker for the German-language Wikipedia. He downloaded approximately 30 GB of text files, parsed them for ISBNs (found in templates and article text), and then ran the checksum algorithm. His initial intent was to identify two categories of errors:

  1. Human mistakes: Inadvertent typos, switched numbers, or other honest errors made by contributors.
  2. Legacy publisher errors: A peculiar issue from the 1980s and 1990s where some publishers, misunderstanding checksums, simply incremented ISBNs by one for subsequent volumes. Libraries developed ways to handle these "technically wrong but functionally accepted" ISBNs, sometimes using templates like German Wikipedia's "ESPN fudge" to acknowledge the error while maintaining functionality.

However, as Schindler delved deeper, he uncovered a third, unanticipated, and far more concerning category of errors. This discovery highlighted a systemic vulnerability that extended beyond simple human or historical publishing errors, pointing towards a new era of information integrity challenges. The problem of circular references, where made-up information in Wikipedia is cited by external sources, which then in turn are used to "verify" the original Wikipedia claim, has existed for years, famously satirized by an XKCD comic. However, the advent of LLMs promised to exacerbate this issue dramatically, creating plausible yet entirely fabricated information at an unprecedented scale. Furthermore, the irony is not lost that Wikipedia, with its rich, multilingual, and structured data, has been a primary source of training data for many of these very LLMs, suggesting a potential for these models to "poison their own pond."

Key Findings

▶ Watch: Wikipedia's philosophy: start research, not end research (3:30)

Mathias Schindler's project to validate ISBNs in German Wikipedia revealed a critical new threat to information integrity: the widespread infiltration of hallucinated literature references generated by Large Language Models (LLMs), primarily ChatGPT. Beyond the anticipated human typos and historical publisher errors, Schindler encountered a third, perplexing category: articles containing multiple literature references that sounded plausible – with familiar-sounding book titles and author names – but which consistently failed checksum validation and could not be located in any library catalog.

This discovery was not an isolated incident. Schindler noted that the University Library of Hogen reported similar issues in 2023, with students requesting librarians to find non-existent literature that matched existing journal titles or issues. Even more disturbingly, the International Committee of the Red Cross issued a strong warning about hallucinated references, as individuals, when informed that their cited literature did not exist, sometimes accused librarians of "hiding the truth." This highlights not just a waste of time for information professionals but also a dangerous erosion of trust.

The core issue, Schindler explained, lies in the fundamental design of LLMs. These models are "machines designed to provide plausible tokens following a certain sequence." Their primary objective is plausibility, not factuality or adherence to real-world truth. This inherent characteristic leads to the generation of "anti-knowledge," which is the antithesis of an encyclopedia's purpose: content that is convincing and contextually appropriate but entirely devoid of factual basis.

The talk also touched upon the ironic feedback loop: Wikipedia's vast, multilingual, and well-structured content has been a prominent source of training data for many LLMs. By re-injecting hallucinated content back into Wikipedia, these models are, in a sense, "poisoning their own pond." Schindler noted that LLM providers are aware of this issue, actively seeking "synthetic-free" content and even paying a premium for data guaranteed to be untainted by AI-generated information.

Within the Wikipedia community, a consensus is emerging: the undeclared use of LLMs for content generation is frowned upon, and individuals have been banned for dumping LLM-created content. The English Wikipedia has even introduced a speedy deletion rule for "blatantly obvious LLM-generated content," though Schindler acknowledged that the bar for detection is constantly being raised as models improve. Beyond ISBN errors, Wikipedians have identified other telltale signs of LLM-generated content, such as:

  • Inconsistent MediaWiki syntax, with models often defaulting to Markdown or using incorrect template names (e.g., English names in German Wikipedia).
  • Inclusion of non-existent parameters within templates.
  • Specific stylistic patterns, such as the overuse of certain adjectives or predictable sentence structures.

Despite the challenges, Schindler's checksum tool demonstrated a tangible impact. German Wikipedia experienced a measurable drop in ISBN mismatches, making it the only Wikipedia edition to show such an improvement, underscoring the effectiveness of targeted human-scale intervention in combating this problem. However, he cautioned that this specific method is unlikely to remain effective indefinitely, as LLMs will eventually learn to generate correct ISBN checksums.

Technical Deep Dive

▶ Watch: Developing an ISBN checksum checker for Wikipedia (6:00)

The technical foundation of Schindler's discovery lies in the robust structure of the International Standard Book Number (ISBN). An ISBN is more than just a sequential identifier; it incorporates an error-detection mechanism in its final digit, known as a checksum. This checksum is crucial because it allows for the verification of an ISBN's validity. For ISBN-10 and ISBN-13 formats, different formulas are used, but both boil down to a weighted sum of the preceding digits. For example, in ISBN-10, each digit is multiplied by a weight from 10 down to 1, and the sum modulo 11 should be 0. ISBN-13 uses a similar concept but with alternating weights of 1 and 3, and a modulo 10 sum. A key mathematical property of these checksums is that they are designed to detect common transcription errors, such as a single digit being mistyped or two adjacent digits being transposed, ensuring that such mistakes will result in a mismatched checksum.

Schindler's ISBN checksum checker was a straightforward but effective tool. He developed a parser (initially in Python, later refined with C code, notably with the assistance of an LLM itself – Cloud Opus, a detail he revealed later in the talk) to process the entire content of the German-language Wikipedia. This involved handling approximately 30 GB of text files, extracting ISBNs found in various locations, including specialized templates and raw article text. Once extracted, the checksum algorithm was applied to each ISBN to verify its mathematical validity. Any ISBN that failed this check was flagged as an error.

The discovery of hallucinated references highlighted a fundamental aspect of Large Language Model (LLM) operation: they are "machines designed to provide plausible tokens following a certain sequence." Unlike traditional databases or search engines, LLMs do not inherently "know" facts or verify information against a truth database. Instead, they generate text based on statistical probabilities of word sequences learned from their vast training data. When prompted to create an article or add references, an LLM will generate plausible-sounding titles, authors, and even ISBN-like numbers because these sequences frequently appear together in its training data. However, without a mechanism to ground these generated identifiers in reality, the checksums and the existence of the referenced material become statistically likely to be incorrect or non-existent. This results in "anti-knowledge" – content that is structurally convincing but factually false.

Current detection methods for LLM-generated content, as identified by Schindler and the Wikipedia community, extend beyond ISBN checksums:

  • Syntax Inconsistencies: LLMs often struggle to consistently produce the specific MediaWiki syntax used by Wikipedia. They might revert to Markdown language, use English template names in non-English Wikipedias, or even invent non-existent parameters for templates. These deviations serve as "telltale signs" for experienced Wikipedians.
  • Stylistic Anomalies: While not a fixed rule, some LLM-generated content exhibits characteristic linguistic patterns, such as an overuse of certain adjectives or a generic, uninspired tone.
  • User Behavior: When confronted, users who have dumped LLM-generated content often exhibit a "reluctance of admitting the use of LLMs" and are unable to provide the prompts they used, which Schindler uses as a "litmus test" for sincerity.

Schindler acknowledged that the ISBN checker, while effective now (leading to a drop in mismatches in German Wikipedia), is a temporary solution. He "strongly believe[s] that at some point the large language model will be able to hallucinate proper check sums," rendering this specific detection method obsolete. This underscores the continuous arms race between generative AI capabilities and defensive measures, requiring constant adaptation and the development of new, more sophisticated detection strategies.

Demo / Proof of Concept

▶ Watch: Discovery of a puzzling third category of unfindable references (9:00)

While the talk did not feature a live, interactive demonstration in the traditional sense, Schindler's entire presentation revolved around the practical application and results of his ISBN checksum checker, effectively serving as a proof of concept for identifying LLM-generated hallucinations.

The development and deployment of this tool were central to his findings. Schindler detailed how he (with later admitted AI assistance) wrote the code, initially in Python and subsequently refined in C, to parse the substantial 30 GB content of the German-language Wikipedia. This parser was designed to extract all ISBNs, whether embedded within structured templates or appearing as raw text in articles. The core functionality involved applying the mathematical checksum algorithm to each extracted ISBN to determine its validity.

The success of this tool was evident in its impact. Schindler explicitly stated that his small checksum tool "contributed to the German language Wikipedia being the only Wikipedia edition [that] had a drop in mismatches in ISBN." This tangible result unequivocally demonstrates the checker's effectiveness in identifying the specific category of errors caused by LLM hallucinations – those plausible-sounding references with mathematically incorrect ISBNs. The tool's ability to pinpoint these otherwise undetectable errors, and the subsequent improvement in data quality, serves as a powerful validation of its concept.

Furthermore, Schindler emphasized transparency by stating that the source code for his checker is released on GitHub and is "openly declared to be a result of large language model use." This not only makes the methodology verifiable but also highlights his personal stance on the ethical use and declaration of AI tools, even in the context of identifying problems created by other LLM applications. The tool's existence and its measurable impact on Wikipedia's data integrity solidify its role as a concrete demonstration of both the problem and a viable, albeit temporary, solution.

Defensive Implications

▶ Watch: The 'Aha!' moment: literature references hallucinated by ChatGPT (10:00)

The proliferation of AI-generated content, particularly the hallucinated references identified by Schindler, necessitates a multi-faceted defensive strategy for content platforms, LLM developers, and the broader information ecosystem.

For Wikipedians and Content Platforms:

  • Implement Robust Validation Tools: Systems should integrate checksum validation for all structured identifiers like ISBNs. This "human-scale" intervention, as demonstrated by Schindler's tool, can significantly improve data quality in specific areas. His plan to integrate actual bibliographic databases from national libraries into future checks will enhance the ability to find more hallucinated content, even if it has a correct ISBN or no ISBN at all.
  • Develop Advanced Detection Mechanisms: Beyond checksums, platforms should invest in tools that can detect common LLM-generated syntax errors (e.g., incorrect MediaWiki templates, Markdown usage, non-existent parameters) and subtle stylistic anomalies. This requires a continuous learning process as LLMs evolve.
  • Establish Clear Policies and Enforcement: The consensus within the Wikipedia community to frown upon undeclared LLM use and implement bans or speedy deletion rules (like English Wikipedia's policy for "blatantly obvious LLM-generated content") is a crucial step. This requires active moderation and a willingness to remove content that cannot be verified.
  • Promote Open Declaration: Encouraging contributors to openly declare their use of LLMs for content generation is vital for transparency and accountability. Schindler's "litmus test" of asking for the prompt used is a practical approach to gauge a user's sincerity and the nature of their LLM interaction.
  • Prioritize Human Oversight: Schindler stressed that "AI is never going to be able to assume responsibility. This is always the responsibility of a human." This means human editors, librarians, and fact-checkers remain indispensable, acting as the ultimate arbiters of truth and quality. The problem "is definitely not enough" to be handled by individual efforts, requiring a broader community response.
  • Consider "Cut-Off" Dates: In cases where AI generation is suspected and the user cannot provide evidence of non-LLM origin, establishing a "cut-off date" (e.g., November 2022, when ChatGPT became widely available) for deletion might be a pragmatic, albeit harsh, solution to prevent unverified content from persisting.

For LLM Developers and Companies:

  • Assume Responsibility for Limitations: LLM providers have a moral and potentially legal obligation to clearly communicate the limitations of their tools, especially regarding factuality. Schindler questioned whether a "slightly light gray on dark gray" disclaimer about inaccuracy is sufficient.
  • Improve Model Grounding and Factuality: A long-term goal for LLM development should be to improve grounding mechanisms that link generated content to verified external data sources, moving beyond mere plausibility.
  • Support "Synthetic-Free" Data: The demand for content "free from synthetic information" and the willingness to pay a premium for it indicates an awareness among LLM providers. This should translate into proactive measures to prevent the poisoning of training data.

Broader Societal and Regulatory Implications:

  • Public Education: There's a critical need to educate the public about the capabilities and, more importantly, the limitations of LLMs, especially the distinction between plausibility and factuality. Many users, as Schindler suggested, might be in a state of "blissful ignorance."
  • Legal and Ethical Frameworks: Schindler posed questions about whether legal or political pathways, such as the EU AI Act, could "shame companies into improving their tools" and mandate greater transparency or accountability for AI-generated misinformation.
  • Preservation of Encyclopedic Knowledge: The talk underscored a fundamental question: "Is there still a place for an encyclopedia in this world?" The integrity of foundational knowledge sources is under threat, requiring a collective effort to preserve and protect them from sophisticated forms of misinformation.

Key Takeaways

  • LLMs Hallucinate Anti-Knowledge: Large Language Models like ChatGPT can generate plausible-sounding but entirely fabricated information, including non-existent literature references with invalid ISBNs, posing a significant threat to factual integrity on platforms like Wikipedia.
  • Technical Checks Are Effective, For Now: Simple technical validations, such as ISBN checksum verification, can identify some forms of LLM-generated "anti-knowledge." Schindler's tool led to a measurable drop in ISBN mismatches in German Wikipedia, demonstrating the immediate impact of such interventions.
  • Undeclared LLM Use is Problematic: The Wikipedia community is building a consensus against the undeclared use of LLMs for content generation, leading to bans and specific deletion rules, highlighting the ethical concerns and the need for transparency.
  • Human Responsibility is Paramount: AI tools do not assume responsibility; the ultimate accountability for the content generated by LLMs lies with human users and developers. Clear declaration of AI use and understanding of its limitations are critical.
  • The Problem is Evolving: Current detection methods, like ISBN checksums or syntax error identification, are temporary. LLMs will likely adapt to generate valid checksums and more consistent syntax, necessitating continuous innovation in defensive strategies and a reliance on deeper semantic and factual verification.
  • Integrity of Knowledge is at Stake: The infiltration of "anti-knowledge" into foundational sources like Wikipedia threatens to "poison the well" of LLM training data and erode public trust in information, demanding a multi-faceted response involving technical, community, and regulatory efforts.

About the Speaker(s)

Mathias Schindler is a long-standing and deeply committed member of the Wikipedia community. He joined Wikipedia in 2003, merely two years after its founding, and went on to become a co-founder of Wikimedia Germany. His involvement extends to previously being an employee of Wikimedia Germany, contributing significantly to the organizational backbone of the movement. Schindler is a familiar face at major technology conferences, having been a passive attendee at numerous Chaos Communication Congress (CCC) events before stepping onto the stage as a speaker. While he currently works at a tech company that utilizes AI, he explicitly stated that his presentation on AI-generated content in Wikipedia was delivered in a private capacity as a Wikipedian, drawing on his extensive experience and dedication to the project.

All talks from 39th Chaos Communication Congress (39C3): Power Cycles