Launching soon. Preferium for agencies is in private preview — partner registration isn’t open yet. Join the waitlist →

All articles

Publishers weigh a Google exit — and crawler defaults flip on September 15

The WSJ reports USA Today, Politico, Reuters, The Economist and Reddit are all reconsidering Google. Meanwhile Cloudflare changes AI-crawler defaults on September 15 — and multi-purpose bots get caught in between.

Photograph: a wide video wall in a bright operations room showing an access-control board of rows with green and red status pills.

TL;DR: The crawler bargain — we index you, you get traffic — is being renegotiated in public. The Wall Street Journal reported on July 21 that USA Today, Politico, Reuters, The Economist, People Inc. and Reddit are all weighing whether to block Google’s crawler outright, and had to correct its own traffic figures a day later. For agencies, the operational story isn’t the publisher politics: it’s that on September 15, 2026 Cloudflare changes its AI-crawler defaults, and crawlers that do more than one job — Googlebot, Bingbot, Applebot — are governed by the most restrictive rule you have set. Crawler access has quietly become a client setting somebody has to own.

What actually happened

Two separate developments landed in the same news cycle, and they belong together.

Publishers are openly discussing an exit. Per PPC Land’s write-up of the WSJ report, five major publishers plus Reddit are evaluating whether to block Google’s crawler entirely — the blunt version of a complaint agencies have heard all year. Nieman Lab covered the same reporting. Reddit’s position is the sharpest: it signed a $60 million-a-year licensing agreement in 2024 letting Google use its content for AI training, and that contract is up for renewal.

The traffic numbers deserve care, because the first published version was wrong. The Journal corrected its Semrush-based figures one day after publication. The corrected declines: USA Today 28%, Reuters 35%, Business Insider 43%, The Washington Post 44% — while The Guardian and the BBC actually gained 16% and 15%. Same sourcing, per PPC Land. The initial, steeper numbers circulated widely before the correction did; if you pasted them into a client deck last week, go fix the deck. This is the AI-visibility measurement problem in miniature — the number that spreads fastest is rarely the number that survives review.

Google’s counter-position, for the record: SVP Nick Fox said on July 17 that Google now sends “billions of clicks to websites every week through AI features in Search alone,” reported by Search Engine Land. No baseline, no denominator, no methodology published. Treat it as a position, not a finding.

The infrastructure moved underneath all of it. On July 1, Cloudflare announced a new taxonomy for AI traffic, splitting bots by behaviour rather than by company: Search (collecting or indexing content), Agent (acting on a person’s behalf in real time), and Training (collecting content to train or fine-tune a model). Each category can be set independently to allow, block, or block only on ad-bearing pages.

The dated part is the default. Per Cloudflare’s changelog entry, from September 15, 2026 new domains onboarding to Cloudflare get Training and Agent blocked by default on ad-supported pages, while Search stays allowed. Existing customers keep their current settings and can opt out of the new defaults at any point before that date.

The part that will bite: multi-purpose crawlers

Here is the detail worth putting in front of every technical lead you work with.

Googlebot, Bingbot and Applebot are not single-purpose bots. They crawl for search indexing and their output feeds AI systems. Cloudflare is explicit that multi-purpose crawlers are evaluated against all of their behaviours and the most restrictive applicable rule wins — so a site that blocks Training also blocks the crawlers that combine search indexing with training.

Read that twice, because it inverts the usual assumption. “Block AI training, keep search” is not a setting you can express by ticking one box. On Cloudflare’s model, choosing to block Training can take classic search indexing with it. That is precisely the trap the publisher exodus story describes — being forced to choose between search visibility and unlicensed AI training — except now it is a dashboard toggle sitting in front of clients who never read the WSJ.

That is also why Cloudflare tells existing customers to confirm their intent in zone Settings before September 15 rather than after. A default that changes on a date is a deadline whether or not anyone treated it as one.

What agencies should do before September 15

This is unglamorous, checkable work — which is exactly why it gets skipped.

  1. Inventory which client sites sit behind Cloudflare. For most agency portfolios this is a meaningful share, and nobody has a current list.
  2. Read the actual crawler settings per zone, not the settings you assume were inherited. Note anything already blocking Training.
  3. Confirm intent in writing with each client. “Do you want to be indexed by search engines whose crawlers also feed AI systems?” is a business question, not a technical one — and it now has a deadline.
  4. Separate the two decisions in the conversation. Blocking AI training is defensible for a paywalled publisher with licensing leverage. It is rarely defensible for a service business that needs to be found. Our take on the AI Overviews opt-out applies here too: removing yourself from a surface you are losing on does not win it back.
  5. Verify after the date, don’t assume. Fetch key pages as each major bot user-agent and confirm a 200. Silent crawler blocks look exactly like normal traffic in analytics until rankings move weeks later.

The wider point is one this blog keeps circling: AI visibility is decided by machine-readable plumbing far more than by content strategy. Whether a bot can reach the page, and what it receives when it does, is the whole game — the same reason AI crawlers not executing JavaScript quietly wrecks tool-applied optimisations. A crawler-access default flipping on a calendar date is that principle with a due date attached.

Key takeaways

  • The WSJ reported on July 21 that USA Today, Politico, Reuters, The Economist, People Inc. and Reddit are weighing blocking Google’s crawler; Reddit’s $60M/year Google training licence is up for renewal.
  • The traffic figures were corrected a day later — USA Today 28%, Reuters 35%, Business Insider 43%, Washington Post 44%, with The Guardian and BBC up 16% and 15%. Use the corrected set.
  • Cloudflare now classifies bots by behaviour — Search, Agent, Training — each independently set to allow, block, or block on ad-bearing pages.
  • September 15, 2026 is a real deadline. New domains get Training and Agent blocked by default on ad-supported pages; existing customers must opt out before that date if they don’t want the new defaults.
  • Multi-purpose crawlers are governed by the most restrictive rule. Blocking Training can also block Googlebot, Bingbot and Applebot — meaning a well-intentioned AI-training block can cost classic search indexing.
  • Audit crawler settings per client zone now, confirm intent in writing, and verify with real fetches afterwards rather than trusting the dashboard.

Preferium runs 47 technical health checks against client pages and verifies every change in a real browser, so a crawler-access regression shows up as a finding rather than as next quarter’s traffic mystery.

More articles Become a partner