Google Made the Web Worse and Is Now Making AI the Same Way

Photo: Unsplash

Platform Dynamics

Google Made the Web Worse and Is Now Making AI the Same Way

The web optimized for Google search and became worse for humans — AI is following the same optimization path and it leads to the same place
googleseoartificial-intelligenceplatform-decaysearch

In 2008, if you searched Google for a recipe, you got a recipe. By 2018, if you searched Google for a recipe, you got twelve paragraphs about the writer’s grandmother’s relationship with food, a pop-up asking for your email address, three autoplay videos, and eventually a recipe. By 2024, that recipe page was generating ad revenue, was ranking top five for its target keyword, and was read by approximately nobody who remembered it.

This was not a failure of Google’s algorithm. It was the algorithm working exactly as designed — optimizing for signals of content quality that content creators learned to fake. The algorithm wanted engagement signals. The content industry produced engagement-optimized garbage. The algorithm couldn’t tell the difference.

AI is now doing this in real time, at scale.

How SEO Broke the Web

The Google search algorithm, fundamentally, tries to identify which web pages best answer a given query. It does this by analyzing hundreds of signals: incoming links (a page that many other sites link to is probably authoritative), content structure (does the page have clear headers that relate to the query?), user behavior (do people who click this link stay on the page, or immediately hit back?), domain authority, page loading speed, and dozens more.

These signals are proxies for quality. A page that many people link to is probably good. A page people read to completion is probably good. The problem is that once you publish the signals, people optimize the signals rather than the quality. This is Goodhart’s Law operating on the web: the measure became the target.

What followed was an entire industry — search engine optimization — dedicated to manufacturing the signals of quality without the substance. Link farms that exchanged links to inflate domain authority. Content farms producing thousands of pages targeting specific keywords, written to human-legible minimums and stuffed with the target phrase at exactly the keyword density Google rewarded. “Engagement” metrics gamed by posting outrage bait that made people spend ten minutes leaving angry comments, which the algorithm counted as high engagement, which drove traffic, which drove ad revenue.

The web didn’t become worse because Google was evil. It became worse because Google was powerful and transparent. Everyone building content knew Google’s algorithm determined their distribution. They optimized for the algorithm. The algorithm’s proxies for quality stopped tracking quality. The result: the web of 2024 contains vastly more content than the web of 2004 and is substantially harder to extract useful information from.

The New Optimization Target

AI systems — specifically large language models serving as information retrieval interfaces — have introduced a new optimization target that content creators are already learning to hit.

When Google started rolling out AI Overviews in 2024, it changed the fundamental incentive structure for web content in a specific way: instead of trying to rank first in search results (where users see the page title and snippet and click through), content creators now need to get cited by the AI Overview. The AI Overview is what most users read and stop there. Getting cited means the content appears above the results, as the answer, attributed to the source. This is a dramatically better position than any traditional SEO ranking.

How do you get cited? The signals AI systems use to select sources for citation are not identical to Google’s ranking signals, but they overlap substantially: authoritative domain, clear structured content, comprehensive coverage of the topic. The new optimization game has already started. Content agencies are explicitly offering “AI-optimization” services — structuring content to be cited by AI responses rather than just ranked in search.

The problem is identical to SEO, one level up. Previously: optimize for signals that Google associates with quality. Now: optimize for signals that AI associates with quality. The gap between “content optimized to be cited by AI” and “content that is actually good” is at least as large as the gap between “content optimized for search rankings” and “content that is actually good.” Possibly larger, because the optimization surface for AI citation is less well-understood and therefore less efficiently exploited — for now. Give it two years.

The Training Data Feedback Loop

The SEO problem was bad. The AI training data problem has an additional dimension that makes it structurally worse: the content being produced gets fed back into the next generation of AI training data.

Large language models are trained on internet text. Internet text is now being systematically polluted with AI-generated content optimized to appear authoritative and comprehensive — the kind of content AI systems select as good training data. When the next generation of models is trained on this data, they internalize patterns produced by content optimized for previous AI systems. The output of that training is then deployed to create more content that fills the training data for the generation after.

Garbage in, garbage out — except the garbage is self-referential. The content produced by GPT-4-optimized SEO farms becomes training data for GPT-5, which produces slightly different optimized content, which becomes training data for GPT-6. The loop is slow and its effects are hard to isolate, but the direction is clear. Model collapse — the phenomenon where models trained on AI-generated data degrade in quality — has been documented in academic literature. The real-world version is messier but the dynamic is the same.

This is the difference between SEO breaking the web for humans and AI training data feedback loops potentially breaking AI systems’ ability to learn from web text at all.

The Humans Who Did This Deliberately

It’s worth being specific about what I mean by “optimizing for AI,” because some of it is clearly coordinated and deliberate.

In late 2023, multiple SEO firms published playbooks for “AI content optimization” — essentially, instructions for structuring content so that AI systems identify it as authoritative and cite it in responses. The advice includes: use structured data markup (schema.org), write in FAQ format with explicit question-and-answer structure, include precise statistics with citations (even when those statistics are invented, since AI systems cannot reliably distinguish accurate from fabricated data with proper citation formatting), and produce extremely comprehensive coverage of topics in a way that superficially mimics what a thorough expert might write.

Some of this advice is unobjectionable — structured data markup is genuinely helpful. Some of it is actively pernicious. “Include precise statistics” becomes “invent precise-sounding statistics,” because the incentive is citation, not accuracy. AI systems are notably bad at verifying the accuracy of specific numerical claims within documents that otherwise appear authoritative. This is an exploitable vulnerability, and it is being exploited.

There’s a direct line from keyword-stuffed recipe blogs to AI-cited Wikipedia-adjacent content farms producing plausible-sounding health misinformation at industrial scale. The mechanism is identical. The consequences are substantially worse.

What Would Have to Change

Breaking the SEO-for-AI cycle requires changing the fundamental incentive structure of content creation on the web. This is much harder than it sounds.

Google tried to improve search quality through algorithmic updates — the Panda update (2011) targeted content farms, the Penguin update (2012) targeted link manipulation, the Helpful Content update (2022) explicitly targeted “content made for search engines rather than people.” These updates are real and they matter. They also inevitably get circumvented, because the economic incentive to circumvent them is enormous and the people doing the circumventing are sophisticated, well-resourced, and operating in a game-theoretic environment where staying ahead of the algorithm is a viable profession.

For AI to avoid the same fate, you’d need evaluation systems that can distinguish content quality from content optimized to appear high-quality — and you’d need those systems to remain ahead of the optimizers. This is fundamentally the same problem Google has been losing for twenty years. There’s no particular reason to expect AI companies to solve it faster.

The more honest answer: some amount of this decay is probably inevitable, just as some amount of SEO garbage was inevitable once search became economically important. The question is whether it reaches a level that degrades the usefulness of AI systems materially, or whether it can be contained below the threshold of catastrophic failure.

The web is still functional. It’s much worse than it would have been without search-engine optimization destroying the incentive to write for humans. AI is functional. Whether it’s still functional in ten years, after a decade of AI-optimized content polluting its training pipelines — that’s the question that nobody is answering confidently, because nobody knows.

What they do know is that the mechanism is the same. The outcome will be similar, in direction if not in magnitude. The web optimized for algorithms and became harder for humans to use. AI will create the same optimization pressure. The trajectory is not mysterious. It’s just happening faster.

The Publisher’s Dilemma

What makes this dynamic particularly stubborn is that individual publishers cannot solve it through individual decisions.

If The Atlantic decides to stop producing content formatted for AI citation and returns to writing purely for human readers, what happens? Its AI citation rate drops. Its discoverability via AI interfaces drops. The traffic that increasingly flows through AI-mediated search goes elsewhere — to publishers who have continued optimizing for AI citation. The Atlantic loses audience without gaining any cleaner information ecosystem, because one publisher declining to optimize doesn’t change the incentives for the other thousand.

This is a classic collective action problem, identical in structure to the race to the bottom in keyword optimization that produced SEO garbage in the 2010s. Every individual publisher has an incentive to optimize for the dominant distribution channel. The collective outcome of everyone following that incentive is a worse information ecosystem. Nobody is doing anything irrational. The outcome is nonetheless bad.

The 2010s version of this problem had a partial solution: Google periodically updated its algorithm to penalize the most egregious optimization strategies, which preserved enough of the quality signal to prevent complete collapse. Whether AI companies will play the same role — updating their citation selection mechanisms to penalize content obviously optimized for AI citation rather than human readers — depends on whether they have an incentive to do so. That incentive exists only if the quality of AI-cited content directly affects their product quality in ways users can detect and punish by switching to competitors. The feedback loop is slow and indirect.

The Specific Irony

Google’s AI Overviews, in particular, complete the circuit in a way that has a black-comedy quality to it.

Google built the algorithm that incentivized SEO garbage. The SEO garbage degraded web content quality. Google’s AI Overviews now summarize web content — including SEO garbage — and present it as the answer. The AI-generated summaries of AI-optimized content then train future AI models. The company that broke the web is now building a product that accelerates the degradation of the resource its product depends on.

It’s not malice. It’s an emergent outcome of market structure and incentive design. That’s the thing about these platform dynamics: they don’t require anyone to be doing anything wrong. They just require powerful optimization pressure and misaligned incentives. Both conditions are met. The rest follows.

Get the next live webinar in your inbox

One email a month: the upcoming live event + free recording access for subscribers. No spam, unsubscribe anytime.