How Much of the Internet Is Fake Now? Experts Explain the Growing AI Content Problem

Recent studies reveal that a growing share of web pages, product reviews, and news stories are AI-generated or assisted

How Much of the Internet Is AI-Generated?
Generative AI and web bots now drive a vast portion of online activity, making it harder than ever to distinguish authentic content from synthetic material ChatGPT AI-Generated

Search crawlers, advertising infrastructure, social media bots and automated tools have driven online activity for decades. Yet generative artificial intelligence is fundamentally altering the scale of machine-created content, making it increasingly difficult to discern whether a webpage, product review, image or news article was authored by a human at all.

There is no single, definitive metric indicating precisely how much of the web is currently synthetic. Different studies measure distinctly different phenomena, ranging from automated site traffic to AI-generated web pages. However, recent empirical research paints a clearer picture of how rapidly synthetic content is spreading.

A 2026 study conducted by researchers at Imperial College London, Stanford University and the Internet Archive found that approximately 35% of newly published websites were identified as AI-generated or AI-assisted by mid-2025. That figure rose from virtually zero before ChatGPT's public rollout in late 2022.

Separate data from SEO firm Ahrefs points to an even broader footprint. After examining 900,000 newly indexed English-language webpages published in April 2025, its researchers calculated that 74.2% incorporated some degree of AI-generated content.

Crucially, however, only a small minority of these pages were categorised as entirely machine-written; the vast majority comprised a hybrid of human and AI composition.

These figures should not be misinterpreted as proof that 35%, 74% or any other proportion of the broader web is entirely fake. They evaluate specific cross-sections of newly published material. Nevertheless, they highlight just how swiftly AI-assisted publishing has embedded itself across the digital landscape.

The Internet Is Already Full of Bots

Another metric makes the internet's human presence appear even sparser.

According to the 2025 Imperva Bad Bot Report, automated traffic accounted for 51% of total web traffic in 2024, surpassing human activity for the first time in 10 years. The report estimated that malicious bots generated 37% of all web traffic, while legitimate bots, such as search engine crawlers, accounted for the remaining 14%.

This does not imply that 51% of the material people read online is synthetic. Automated traffic includes essential operational activities, such as search indexers and data-retrieval systems. Nor are all bots powered by generative AI models.

This distinction is vital because digital 'traffic' and digital 'content' are fundamentally different concepts.

What these data points ultimately confirm is that machines already dominate web traffic. Generative AI is now introducing a second dimension: automation is no longer confined to visiting websites, but extends to producing articles, consumer reviews, imagery, video and whole domains at a scale that was previously economically unviable.

What Is 'AI Slop'?

The term 'AI slop' has emerged as a catch-all descriptor for vast volumes of low-quality, cheaply produced synthetic material.

This includes generic, SEO-driven articles designed primarily to capture search traffic, synthetic images created to solicit clicks, auto-generated videos and fabricated social media posts. While some of this material is benign or merely poorly executed, other content deliberately misinforms.

Notably, recent academic findings offer a nuanced counterweight to the more alarmist commentary surrounding AI slop.

The collaborative study by Imperial College London, Stanford and the Internet Archive observed that AI-generated and AI-assisted websites correlated with lower semantic diversity, meaning online material became increasingly uniform in phrasing and meaning.

Researchers also noted a higher baseline of positive sentiment in synthetic prose. However, they uncovered no statistically significant evidence that the proliferation of AI-generated text was actively eroding factual accuracy or stylistic variation.

That distinction is paramount. The internet may be growing more formulaic and synthetic without every machine-written article necessarily being false.

The core challenge, therefore, is not simply that AI spreads falsehoods; it is that the sheer volume of synthetic media makes it significantly harder for users to isolate genuinely valuable information from material manufactured solely because it is cheap and effortless to produce.

Fake Reviews Are Becoming Easier to Produce

E-commerce reviews represent another domain where generative AI poses a serious threat to consumer trust. A 2025 joint academic study evaluated whether individuals could distinguish genuine consumer reviews from fake ones generated by large language models.

Across three separate trials, participants recorded an average accuracy rate of just 50.8% — a performance indistinguishable from random guessing. Furthermore, the researchers observed that AI detection systems themselves struggled to differentiate authentic reviews from machine-fabricated ones.

Consequently, an AI system does not need to produce an overtly ridiculous evaluation to effectively manipulate a market. A convincing paragraph detailing a restaurant, hotel or product can be generated in seconds, enabling bad actors to manufacture synthetic consumer experiences en masse.

Regulatory bodies are already framing AI-generated reviews as an explicit consumer-protection priority.

In August 2024, the US Federal Trade Commission finalised a rule prohibiting fake reviews and testimonials, explicitly targeting synthetic submissions. The measure bans businesses from creating, purchasing or distributing deceptive reviews that misrepresent the reviewer's identity or first-hand experience with a product or service.

Additionally, the FTC advises consumers to cross-reference reviews across multiple platforms, remain sceptical of sudden spikes in positive feedback and scrutinise whether endorsements appear genuinely independent or commercial in nature.

AI Can Manufacture Entire Websites

Perhaps more alarming than an isolated machine-written article is the proliferation of entire domains engineered from inception to masquerade as legitimate digital publications.

In March 2026, the cybersecurity firm DoubleVerify revealed it had uncovered a coordinated network exceeding 200 websites dedicated to publishing AI-generated text and media. The firm designated the operation 'AutoBait'.

While these platforms presented themselves as lifestyle publications, their primary function was delivering ad impressions at scale. DoubleVerify reported that the network amassed tens of millions of views, noting that the underlying programme-generation scripts were inadvertently exposed within the sites' JavaScript code.

This case illustrates why fully automated websites differ fundamentally from an individual blogger using ChatGPT to assist with drafting content.

The underlying technology exponentially lowers the cost of publishing hundreds or thousands of pages. This dynamic creates a direct financial incentive to flood the web with content designed not to inform readers, but to capture search queries, display ad impressions and farm clicks.

Google acknowledges the issue. Its official Search guidelines state that while generative AI can assist in content creation, using automation to churn out high volumes of pages without providing distinct value constitutes 'scaled content abuse' and violates search policy.

Fake News Is a Bigger Problem

AI-generated misinformation presents a unique challenge because fabricated narratives can now be packaged in highly persuasive formats.

Researchers have systematically documented the integration of synthetic content across disinformation channels. A 2023 study analysing more than 15 million articles from 3,074 mainstream and fringe websites identified a sharp rise in machine-generated news across both categories, with the most pronounced surge occurring among known misinformation outlets.

Furthermore, generative tools now produce photorealistic imagery, voice clones and synthetic personas. Misinformation no longer arrives in the form of poorly constructed, obviously illegitimate blogs.

A fabricated quotation can now be paired with an AI-generated photograph; a false press release can be spoken via a cloned voice profile; a non-existent expert can be equipped with a polished biography and active social media footprint.

This shift creates an information ecosystem where individual pieces of media require proactive verification, rather than automatic trust based on polished visual presentation.

Google AI Overview Adds Another Layer

The surge in synthetic content takes on added significance as the search experience itself relies more heavily on artificial intelligence.

Google's AI Overview feature serves AI-generated summaries at the top of select search queries, alongside citation links to reference sites. Google explicitly notes that these automated summaries can contain errors.

This setup introduces a potential feedback loop.

As low-quality, synthetic web pages populate the internet, search indexers must evaluate which sources are credible. If an automated model subsequently synthesises answers from these indexed pages, inaccuracies or weak sources risk being integrated directly into the primary summary presented to users.

Research published in 2026 examining Google Search, Gemini and AI Overviews found that AI Overviews appeared for 51.5% of analysed real-user queries. The study also revealed that the source material cited within AI Overviews frequently diverged from the ranking results yielded by standard Google Search queries.

This does not mean AI Overviews are inherently inaccurate, but it highlights that an automated summary should not be mistaken for primary source material.

When verifying critical claims, inspecting the foundational source remains indispensable.

How to Tell What Is Real

There is no definitive visual test to identify machine-generated content.

Automated detection software remains prone to false positives, and traditional indicators — such as awkward phrasing or visual artefacts — are fading as generative models mature. Consequently, readers must focus on information provenance: identifying where content originated and verifying whether it can be corroborated independently.

For news coverage, confirm whether established journalism outlets are reporting the story and whether key claims are attributed to named, verifiable sources. For consumer reviews, prioritise verified purchasers, detailed personal accounts and consistent sentiment across independent platforms rather than relying on an isolated rating score.

When evaluating unfamiliar websites, examine who operates the platform, confirm whether author credentials are listed, assess whether articles feature original reporting and check whether the domain regularly outputs unusually high volumes of generic content.

Visual and auditory media demand equal scrutiny. A compelling photograph is no longer proof of a real event, just as an authentic-sounding audio file is no longer definitive proof of a recorded statement.

Finally, when relying on Google AI Overviews, treat the output as a preliminary orientation rather than a final verdict. Always inspect the underlying citations, particularly when evaluating medical, financial, legal or breaking-news developments.

The Internet Is Not Dead — But Trust Is Changing

The premise that the internet has transformed into a total fiction is unsupported by empirical evidence. Humans continue to generate vast quantities of online material, and AI-assisted content is not inherently deceptive or devoid of value.

The reality is more nuanced.

A growing percentage of new online material is authored with machine assistance, while automated traffic already accounts for more than half of overall web activity. Concurrently, generative tools have lowered the cost barrier to fabricating product reviews, synthetic news and low-value web networks.

The defining crisis of the modern web is not that digital content has become entirely artificial. It is that the economic cost of engineering something that appears authentic has collapsed.

For readers, this dynamic fundamentally reshapes how we navigate digital spaces. A visually polished webpage, a thorough product review, a crisp photograph or an authoritative search summary can no longer serve as standalone proof of authenticity.

The internet remains rich with human perspective and genuine expertise. However, accessing it increasingly demands something AI can synthesise cheaply, but never replace: deliberate human discernment.


Frequently Asked Questions

  • What is AI-generated content?
    AI-generated content refers to material created by artificial intelligence systems, such as text, images, or videos, often with minimal human intervention.
  • How does AI-generated content affect the internet?
    AI-generated content increases the volume of synthetic material online, making it harder to discern genuine information from machine-produced content.
  • What is 'AI slop'?
    'AI slop' is a term used to describe low-quality, cheaply produced synthetic material generated by AI, often for SEO or clickbait purposes.
  • How can consumers identify fake reviews?
    Consumers can identify fake reviews by cross-referencing reviews across multiple platforms, being skeptical of sudden spikes in positive feedback, and checking for genuine independence in endorsements.
  • What is the impact of AI on web traffic?
    AI has increased automated web traffic, with bots accounting for a significant portion of total web activity, surpassing human activity in recent years.