Written by Xiao Bing
The Pew Research Center released a large-scale analysis this week based on the Common Crawl web archive. The research team randomly sampled 10,000 English-language web pages from each of 49 web crawls between January 2021 and July 2026, totaling approximately 490,000 pages, and then scanned each page individually using an AI detection model called Open Pangram.
The conclusion can be summarized in two numbers. In the full snapshot of July 2026, approximately 10% of English-language web pages showed "significant signs of AI writing or deep editing." For web pages launched after the release of ChatGPT (November 30, 2022), this proportion surged to 35%.
The 35% figure does not mean that one-third of the internet is written by AI, as the web contains a vast amount of content created before 2022; however, it points to a clear trend: AI involvement in new content is transitioning from an exception to the norm.
Where is AI most prevalent?
The distribution is uneven.
On .com domains, about one in ten pages exhibited AI-generated characteristics in the 2026 sample—more than double the rate on .org domains (4.6%) and ten times the rate on .edu and .gov domains (both around 1%), indicating that the commercial internet is adopting AI far faster than academic and government institutions.
Pew's research also dissected the "linguistic fingerprint" of AI-generated writing.
Compared to 2023, the 2026新版网页 shows a doubling in the use of em dashes, a 63% increase in Oxford commas, more than a doubling in AI-preferred vocabulary (such as delve, interplay, tapestry, pivotal), and a near tripling in the frequency of negative parallel constructions like "it's not just X, it's Y." Individually, these changes do not constitute conclusive evidence, but when aggregated statistically, they point to a clear signal: the linguistic texture of the new web pages is trending toward the characteristic distribution of AI-generated text.
The bigger picture: Who is the audience?
Pew's report focuses on "who is writing," but when viewed alongside another set of data, the picture becomes more complete.
On June 3, 2026, Cloudflare CEO Matthew Prince announced on social media a milestone he hadn’t expected to arrive so soon: AI bot traffic surpassed human traffic for the first time, accounting for 57.4% of global web HTTP requests.
He originally anticipated this crossover point would occur by the end of 2027, but it arrived 18 months early. HUMAN Security’s 2026 report provided data in the same direction: throughout 2025, AI-driven automated traffic grew eight times faster than human traffic, with “agentic AI” (AI that performs tasks on behalf of users) seeing year-over-year traffic growth of nearly 8,000%.
Put these two things together: On the supply side, one-third of newly created web pages involve AI in their writing; on the demand side, more than half of the "readers" are machines. An infrastructure originally built for human communication is rapidly becoming an information pipeline written by machines for machines.
Self-contamination cycle
This is not just a story about declining content quality—it touches on fundamental issues within the AI industry itself.
In 2024, Nature published a paper by researchers from Oxford and Cambridge demonstrating that AI models undergoing recursive training (training the next generation of AI on data generated by AI) experience "model collapse": outputs gradually diverge from the true data distribution, lose rare patterns at the tails of the distribution, and after several iterations generate increasingly homogeneous and even meaningless content. Researchers at Epoch AI predict that high-quality, human-generated text suitable for training AI may be exhausted between 2026 and 2032.
The cycle is already underway: AI models read the human internet and batch-generate new web pages; these pages are indexed by search engines and archived by crawlers like Common Crawl; the next generation of AI models is then trained on this data. With each cycle, the proportion of "AI-written-by-AI" content in the training data increases, while signals from genuine human first-hand experience become increasingly diluted.
If the signal-to-noise ratio of the internet continues to deteriorate, content creators reliant on search traffic will be the first affected.
Search results are flooded with homogenized AI rewritten content, making it increasingly difficult for readers to distinguish between "originals" and "copies." Google has already begun replacing some search result links with AI Overviews, and another Pew study from July found that only 20% of users consider AI search summaries "very useful," with just 6% saying they "very much trust" them.
When AI-generated paraphrases are everywhere, media outlets that provide the following will command a scarcity premium: firsthand information from the scene, exclusive data and documents, perspectives and judgments traceable to specific human sources, and editorial taste that readers are willing to pay for.
From another perspective, the metrics used to measure a media outlet’s competitiveness are quietly changing.
The number of articles published is no longer a barrier; AI can generate a thousand articles in a day. The real barrier is "provable human originality": Did a journalist uncover the information in this article, or did the model fabricate it? Who made this judgment—an editor with industry experience, or a patched-together prompt?
In an internet where 35% of new web pages show signs of AI generation, being able to answer the above questions well is a competitive advantage.
