Stanford Study: 35% of New Websites Will Be AI-Generated by Mid-2025

icon币界网
Share
AI summary iconSummary
A new study by Stanford, Imperial College London, and the Internet Archive, cited by Bitjie, shows that 35% of new websites by mid-2025 are AI-generated or AI-assisted. Researchers analyzed 33 months of website data using an AI text detector. AI-generated sites exhibited higher semantic similarity and positive sentiment but no decline in factual accuracy. The findings emerge amid growing interest in AI and crypto news, as well as new token listings.
CoinDesk reports:

A new study reveals that by mid-2025, 35% of content on the internet will be generated by artificial intelligence. This research, conducted jointly by Stanford University, Imperial College London, and the Internet Archive, shows that 35% of newly launched websites will be classified as AI-generated or AI-assisted. Before the launch of ChatGPT in November 2022, this percentage was nearly zero.

“The speed at which artificial intelligence has taken over the web has shocked me,” said Jonáš Doležal, researcher at Imperial College London and co-author of the paper, telling 404 Media. “After decades shaped by humans, a large part of the internet has been defined by AI in just three years.”

This study of the article titled "The Impact of AI-Generated Text on the Internet" used website snapshot data from the Internet Archive spanning 33 months. Internet Archive and classified each page using the AI text detector named Pangram v3.

Confirmed hazard: atmosphere, not facts

Researchers tested six hypotheses regarding the impact of AI-generated content on the internet, but only two hypotheses withstood data scrutiny.

First point: We are becoming a group of identical, foolish NPCs... or, more scientifically, semantic diversity on the network is decreasing.

Websites generated by artificial intelligence are 33% higher in semantic similarity than those written by humans. The same ideas are always expressed in nearly identical ways.

The paper points out that the online Overton window may be narrowing, but not due to censorship or coordinated action, rather because language models optimize their outputs to align with the training distribution.

Second point: The network is becoming increasingly fun.

The positive sentiment score of AI-generated content is more than 107% higher than that of human-generated content. Researchers link this to the well-documented flattery tendency of LLMs (language learning models)—trained on human-approved signals, their generated text feels sanitized, fluid, and consistently upbeat.

The internet is flooded with cheerful, homogenized content that could大规模 marginalize human dissent even without any action being taken.

Although the public generally believes that AI-generated content reduces factual accuracy on the internet, this study found no statistically significant evidence to support this claim. Researchers also found no significant correlation between the prevalence of AI and the rate of factual errors.

This monocultural assumption—that AI would flatten individual voices into a uniform, generic vocal range—is the most widely agreed-upon viewpoint among respondents (83% agreed). However, the data does not substantiate this assumption. Analysis at the character level found that the adoption of AI has not led to a statistically significant increase in stylistic homogenization.

The model collapse issue has truly arrived.

The broader stakes extend far beyond话语 quality. At an AI adoption rate of 35%, the theoretical risk is... model collapse—a decline in performance of future models trained on AI-generated data—which has shifted from academic speculation to real-world concern. Future foundational models trained on contemporary web crawlers will inevitably absorb vast amounts of AI-generated data that are notably lacking in semantic diversity.

The team is currently collaborating with the Internet Archive to turn this research into an ongoing, real-time monitoring tool that tracks the share of artificial intelligence online in real time, rather than as a one-time snapshot.

A U.S. survey conducted alongside this study found that most Americans already believe all six negative assumptions, including those unsupported by data. People who rarely use AI are 12% more likely to believe in AI’s harms than those who use it frequently. Internet Death Theory believers, take note: the internet hasn’t died, but 35% of newly added content may be some form of zombie content.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.