Technical SEOGoogle
Internal search pages still need deliberate crawl controls
Google removed its old policy wording on internal search results. Review crawl waste and unwanted indexed pages before changing the site's existing controls.

Jump to a section
Google’s removal of an old policy reference to internal search pages does not settle whether yours should be crawled. The useful question is what those pages contain, how many URLs the search form can generate and whether any are already indexed.
The URLs a site search can generate
Your site’s search pages. Not Google’s results — your own. Someone types “blue jacket” into your shop’s search box, your site builds a results page for that phrase, and that page has an address, usually something like yoursite.com/search?q=blue+jacket. Type something else and you get a different address. There’s no limit to it. Every phrase anyone could type produces another page.
robots.txt. A small text file at the root of a website, listing instructions for the automated programs that read the web. It tells them which parts to skip. A line reading Disallow: /search means “stay out of there”. That line is what the old rule was about.
The question in this story is whether you should let Google wander into the endless supply of pages your own search box can produce.
What changed
Since Google’s original 2007 guidelines, “block your internal search results” sat in the rulebook as near enough a requirement. It was the kind of item that turns up on a technical audit checklist with an official citation beside it. It isn’t there any more. Mueller’s words, via PPC Land’s write-up: “Nowadays, we don’t have that listed in the search policies.”
Google didn’t say these pages are now welcome. It didn’t change how anything works, and it didn’t withdraw the technical guidance explaining why the pages caused trouble. What was probably always practical advice has stopped being written down as a rule.
Why you should still block them
Mueller gave two reasons in the same episode. Both are about how the machinery behaves, not about what the rulebook says.
It wastes Google’s attention on your site. Google doesn’t read every page on the internet. Its guidance on crawl budget calls the web “a nearly infinite space, exceeding Google’s ability to explore every publicly accessible URL”, and tells site owners to consolidate duplicates so crawling goes to unique content rather than unique addresses. Your search box can produce addresses without end. Time Google spends working through them is time it doesn’t spend on your product and article pages. Mueller put it plainly: “It’s just like, you’re being very inefficient.”
Strangers can put words on your site. A search page displays whatever was typed into it. That has long been a route for getting spam onto a respectable domain and from there into search results — gambling, pharmaceuticals, adult content. Mueller noted it can get a site flagged as hacked in Search Console no matter what the guidelines say.
Check the rules for other crawlers too
Google’s crawler is well-behaved. It has a documented limit on how much it will fetch, and Google has every reason to be efficient about it. The crawlers now gathering material for AI answers are a more mixed population, and here’s the catch: an instruction addressed to Google only applies to Google.
If your robots.txt blocks the search pages under a heading that names Googlebot specifically, rather than one meaning “everyone”, then GPTBot, ClaudeBot, PerplexityBot and whatever launches next can walk straight into the space Google itself declines to enter.
That’s a server-load problem, at a moment when crawler defaults across the web are being rewritten. It’s also a quality problem. An AI system reading your search pages is learning about your business from thin, repetitive, auto-generated copies of your catalogue instead of the pages you actually wrote. Inspect the response those crawlers receive, rather than assuming it matches the browser view.
What to actually do
For most sites, nothing. Check it, then move on.
- Confirm the search pages are still blocked, and check who the instruction is addressed to. Check the applicable user-agent groups and how each relevant crawler interprets them. A robots.txt instruction is not an enforced access block.
- Choose the control for the problem. robots.txt can reduce unwanted crawling. If a URL is already indexed, blocking crawling can prevent Google from seeing a
noindexdirective. Plan index removal separately; Google’s crawl-budget guidance explains the distinction. - Don’t delete a blocking rule nobody can explain. If no one remembers why it’s there, write down why, don’t remove it.
- Update your audit template, not your website. If your checklist cites this as an official Google requirement, that citation is now wrong even though the advice still holds. Fix the wording before a client finds it.
- Don’t confuse this with proper category pages. A curated listing page with a stable address is a different animal from an open search box, and can deserve to be in Google on its own merits. This rule was only ever about the search box.


