The Internet of Probability
Introduction
Large language models (LLMs) are transforming the internet, from a repository of knowledge into an engine of plausibility. When AI models came online for general use, the internet was already existing in a post-truth reality. Misinformation and disinformation abounded, unfettered, influenced by the forceful contributions of social media platforms and aggressive marketing platforms. Given the pre-conditions upon which AI systems came online, the internet was primed for the entry of questionable content, where anyone’s best guess could verify proclaimed facts. And in turn, consumers of information online, naturally took to AI systems, as we were acclimated to our echo chambers and oddball claims containing questionable facts.
Between 2018 and 2021, there was a segue between the rise of echo chambers and the transition to AI-infused realities, where swarms of humans backed up beliefs and claims with the retort, “do your research”, putting the onus back on the one questioning the shaky claims1
Conspiracy started outweighing truth and any truth could be reconstructed through research, often pointing to unreliable sources and a whole lot of reading between the lines. This transition period set the stage for what was to come, where claims could exist without provenance nor substantiated by qualified research. And so when commercial AI models were made available to the general consumer, the large majority of consumers accepted knowledge without provenance and research without verified citations.
The Advent of Information Generation
A population trained for years to navigate a web of unverifiable claims, is the same audience primed to appreciate and accept AI-generated output. The interface was frictionless because the epistemological habits were already there. An LLM communicated in the same declarative tone as a viral post or a confident comment thread.
Fine-tuned to the constructs of marketing copy, LLMs offer information resolution where the information landscape, comparatively, offers noise. Its answers are clean and for the most part, complete, while the fractured web have made it increasingly difficult to find reliable answers and information.
For lack of transparency, guardrails and policy, it is easy to treat AI as an information knowledge retrieval tool, as humans had not experienced the expeditious production of information and knowledge within any prior information system. The experience of receiving an answer felt indistinguishable from the experience of finding one. Which leads us to today, where information retrieval is being replaced by information generation, which means that the act of retrieval across multiple sources is fast being replaced by single source information delivery systems.
Fluent Guessing
When you ask an AI system a question, it does not retrieve truth—it computes probability. It calculates, across billions of learned associations, which token — which word, which phrase, which claim — are statistically most likely to follow the ones that preceded it. The output resembles knowledge and sounds authoritative. It arrives in confident, well-formed prose. But at its computational core, it is a sophisticated act of pattern completion and rarely a declaration of fact grounded in provenance. As researchers at Oxford’s Internet Institute put it with arresting plainness: LLM outputs represent a “happy statistical accident” — apparent truthfulness that cannot be relied upon because it was never engineered to be true.
The probabilistic architecture of large language models is an epistemological event — one whose consequences are amplified by, and structurally continuous with, the post-truth information environment that social media has been constructing for over a decade.
The internet is in the process of becoming a probability machine: a vast, self-reinforcing ecosystem in which the plausible stands in for the true, provenance has been stripped from knowledge, and the conditions for distinguishing fact from fiction are disappearing at scale. Understanding how we arrived here — and what is at stake — requires holding the technical and the social dimensions of this crisis together, because neither can be understood without the other.
Probability as Architecture
To understand the stakes, we must begin with the mechanics. Large language models are trained on massive corpora of text — drawn primarily from the open internet — and learn the statistical relationships between words, phrases, and concepts across that data. They are designed to predict which token comes next in any given sequence. This is its foundational operation.
As Mark Coeckelbergh, Professor of Philosophy at the University of Vienna, summarizes in his systematic analysis of LLMs and democratic truth: these systems have “been trained to optimize the goal of producing one plausible-sounding word after another rather than actually engage with the meaning of language.”2
This is the industrial distributional hypothesis. The theoretical core of LLMs is grounded in the idea that words and meanings can be characterized by their statistical co-occurrence patterns — that you shall know a word by the company it keeps.
As Philip Resnik of the University of Maryland argues in his landmark position paper published in Computational Linguistics, this is a legitimate and powerful framework for processing language. What it does not do, and cannot do by design, is establish a relationship between language and the world it refers to. An LLM trained to predict likely text has no mechanism for evaluating whether the text it generates corresponds to actual facts and truth. It has learned what truth-claiming language looks like, not what truth is.3
The result is a class of outputs that researchers at Oxford’s Internet Institute have coined “careless speech” — responses that are “plausible, helpful and confident, but that contain factual inaccuracies, misleading references and biased information.” Unlike deliberate misinformation or even hallucination in the clinical sense, careless speech is systemic: it is the predictable output of systems optimized for fluency without a binding obligation to accuracy.4
The statistical veneer of confidence makes careless speech especially dangerous. As Oxford’s Professor Brent Mittelstadt explains: “People using LLMs often anthropomorphise the technology, where they trust it as a human-like information source. This is, in part, due to the design of LLMs as helpful, human-sounding agents that converse with users and answer seemingly any question with confident sounding, well-written text. The result of this is that users can easily be convinced that responses are accurate even when they have no basis in fact or present a biased or partial version of the truth.”5
And bias is not a correctable anomaly layered on top of an otherwise neutral system because it is structural. Resnik’s argument is unequivocal—harmful biases are “an inevitable consequence arising from the design of any large language model as LLMs are currently formulated.” 6
Because LLMs learn from human text, they reproduce human distributions — including the distributions of error, prejudice, historical distortion, and omission that characterize any large, unsystematic sample of human language. Mitigation techniques like reinforcement learning from human feedback (RLHF) adjust outputs at the margins without touching the pre-trained representations that generate bias in the first place. Navigating that tradeoff, Resnik notes, is occurring “in a sea of uncertainty” in which neither the problem nor the intervention is fully understood.7
A Poisoned Well
The probabilistic nature of LLMs would be concerning under any circumstances. It becomes catastrophically concerning when we examine what they are trained on. The internet — the primary training corpus for essentially all large-scale LLMs — is not a library of vetted knowledge. It is, as Ferentinos and colleagues observe at TechPolicy.Press, “a highly adversarial space, where distinguishing fact from falsehood is incredibly difficult.”8
The adversarial character profiles of the internet as a training corpus is the product of decades of structural incentives. State-backed influence campaigns, commercial content farms, coordinated political manipulation, and inadvertent error have flooded the web with content designed to appear authoritative while being systematically misleading.
Russian state actors have deliberately seeded major datasets with manipulative content targeting democratic discourse. Medical misinformation networks generate content that “vastly outperforms content asserting” scientifically grounded medicine on social media engagement metrics — meaning it is disproportionately represented in training data due to its virality.9
The consequence is what researchers at Oxford recognized in their 2023 study. “LLMs are trained on large datasets of text, usually taken from online sources. These can contain false statements, opinions, and creative writing amongst other types of non-factual information.” An LLM trained on such a corpus learns the rhetorical and structural patterns of misinformation — the grammar of political conspiracy, the structure of convincing but unfounded scientific claims — and it reproduces these patterns probabilistically, without any mechanism for flagging that it is doing so.10
The deeper problem is what Ferentinos et al. call “data voids” — areas where very little credible information exists and adversarial groups work deliberately to fill the gap with misleading content, becoming the most statistically prominent sources on those topics. When LLMs encounter queries in these domains, they draw on whatever statistical signal is strongest — the adversarial content that was engineered to dominate the internet. The AI model works, as designed, faithfully reproducing the probability distributions of the corpus it was trained on. No matter if the content lacks facts, truth or presents adversarial content.11
The feedback loop this creates is one of the most alarming features of AI systems and human knowledge. Errors in LLM outputs contaminate trusted public knowledge repositories — Wikipedia entries, code repositories, educational resources — which then enter the training data of the next generation of models.
Subsequent generations of training data becomes epistemic incest. AI systems trained on AI-generated content, in a recursive loop, that degrades the knowledge commons without any external check or policies enforced to ensure the reliability of AI output. probability machines are rewriting the internet in their own image, making it progressively less reliable as a source of ground truth for the systems that will come after.12




