GPTBot Obeyed My robots.txt. Meta's Crawler Ignored It. Today the Build Deleted It.

63,649 filter requests to zero for GPTBot. 937 to 10,566 for Meta's crawler. Same file, same line, same site, same three weeks. And the line is gone as of this afternoon.

Krisada Eaton 11 min read 9 views

One line in a robots.txt file took OpenAI's GPTBot from 63,649 filter-URL requests to exactly zero. The same line, on the same file, on the same day, did nothing at all to Meta's crawler, which went from 937 of those requests to 10,566.

So the answer to whether AI crawlers respect robots.txt is not yes or no. It is per company, and I can now put numbers on four of them.

The part that stings is the last act. The line went in on September 11. It worked for three weeks. This afternoon, about an hour before I sat down to write this up, the nightly sitemap build regenerated robots.txt on all thirty of those subdomains and removed it.

A crawl fix that lives inside a generated file is not a fix. It is a countdown.

What the Filter Grid Does to a Crawl Session

Start with the thing being protected against, because the scale of it is the reason any of this matters.

bars.bellyupjax.com is a small bar directory for Jacksonville. Its sitemap declares 22 URLs. It has twelve venues, four filters that combine, and a sort control. Every combination a visitor can click becomes another URL a crawler can follow, so 22 declared URLs became 82,913 reachable ones.

GPTBot arrived on September 28, 2026 and made 82,129 requests. Quiet the next day. On September 30 it came back for 79,725 more.

Broken out by hour, those days do not look like traffic. They look like a metronome.

  • September 28, 03:00 to 13:00: between 7,038 and 7,198 requests every hour, eleven hours in a row, then a hard stop.
  • September 30, 09:00 to 19:00: between 7,183 and 7,198 requests every hour, eleven hours in a row, then a hard stop.

7,196 requests in an hour is 1.999 per second. Held within nine requests an hour across twenty-two hours on two separate days. That is a rate limiter running out its clock, not a crawler finding my content fascinating.

Over the six days to October 2, GPTBot made 206,380 requests to that site. 206,356 carried a query string. Twenty-four did not. The four Jacksonville bar guides, written for actual people trying to find a bar, were each read exactly one time.

Filter combinations → Real pages
206,356 requests → 24 requests
82,913 distinct URLs → 17 distinct URLs
99.99 percent of the session → 0.01 percent of the session
1,137 MB served → under 1 MB served

The Atlanta Line, and What It Actually Did

None of that was news to me. I found the same trap on the Atlanta network and wrote it up on September 9. Two days later a one-line rule went into robots.txt on the Atlanta category subdomains:

Disallow: /*?

I pulled it out of a backup archive while writing this, because I wanted the exact wording and the exact date rather than my memory of it. The file is stamped September 11 at 23:30, and it survived in the weekly backups through September 26.

Here is what happened on the Atlanta network either side of that line. Before is August 28 to September 11, the full raw retention window up to the change. After is September 12 to October 2.

GPTBot: 63,649 filter requests before, zero after. Not a reduction. Zero, across twenty-one days.

OAI-SearchBot: 19 before, zero after.

Amazonbot: 545 before, 45 after. Down 92 percent, so it mostly honours the rule and not entirely.

meta-externalagent: 937 before, 10,566 after.

Read that last one again. Meta's crawler saw the same file every other crawler saw and increased its filter-URL requests by a factor of eleven. Disallow: /*? uses wildcard syntax, which is a widely supported extension rather than part of the original standard, so a crawler that does not implement wildcards would read that line and match nothing. I cannot tell from my logs whether that is indifference or an unimplemented pattern. I can tell you the effect is the same either way.

Four Crawlers, One robots.txt Line, Three Weeks

Complied fully
GPTBot (OpenAI)
63,649 filter requests before the line. Zero after, across 21 days.
Complied fully
OAI-SearchBot (OpenAI)
19 filter requests before. Zero after.
Mostly complied
Amazonbot (Amazon)
545 filter requests before, 45 after. Down 92 percent, not to zero.
Ignored it
meta-externalagent (Meta)
937 filter requests before, 10,566 after. Up elevenfold with the rule in place.

The Same Split Shows Up Where There Was No Rule At All

The Atlanta result is a before-and-after on one site, so it could be a quirk of timing. It is not, because the unprotected networks show the same vendor pattern from a completely different direction.

Across the fifteen Jacksonville, Orlando and St Augustine category subdomains, over the same six days, with Allow: / and no rule of any kind, six AI crawlers split cleanly into two groups.

Enumerated the filter grid: GPTBot at 565,258 requests and 99.9 percent query strings, meta-externalagent at 88,559 and 94.6 percent, Amazonbot at 28,283 and 99.0 percent.

Touched it zero times: ClaudeBot at 528 requests, OAI-SearchBot at 302, PerplexityBot at 2. All three with no query string at all.

My own warehouse labels ClaudeBot a training crawler, exactly like GPTBot, Amazonbot and meta-externalagent. So the split is not training crawlers against search crawlers. ClaudeBot sits on the training side of my own labels and behaved like the search crawlers.

Whatever separates these six, it is the vendor and not the declared job. I have been carrying a working theory in my research ledger that training and retrieval crawlers respond to different site properties. This week says the real line runs somewhere else, and I would rather learn that from my own logs than keep repeating the tidier version.

The Canonical Tags Were Correct the Whole Time

Every filter URL on every one of these sites returns HTTP 200 with a rel="canonical" tag pointing at the unfiltered page. That is the correct treatment for a duplicate filtered view and it is what my own build standards require. Nothing was misconfigured, on any of the thirty subdomains, at any point in this story.

It also never prevented a single request.

That is not a surprise if you read Google's own documentation more carefully than I had. Google's faceted navigation guidance says using rel="canonical" on these URLs "may, over time, decrease the crawl volume" of the non-canonical versions. May. Over time. Decrease. It is listed below the two methods that actually stop the fetch.

The mechanism is obvious once you say it out loud: a crawler has to download the page to read the tag telling it the page does not matter. Canonical is an indexing instruction delivered by a crawl. It is fine at a hundred duplicate views and it is 206,356 requests at 82,913.

Canonical decides what gets indexed. It has almost nothing to say about what gets crawled. I had those two filed as one thing for longer than I want to admit.

Why Jacksonville Never Got the Line

The fix worked on Atlanta. So why did Jacksonville absorb 206,380 requests three weeks later?

Because the fix was a hand edit to one file on one network, and it was never put into the thing that builds the sites.

Jacksonville, Orlando and St Augustine first appear in my logs on September 18, ten days after I published the Atlanta diagnosis. They launched off the same generator Atlanta came from, which writes a robots.txt containing User-agent: *, Allow: /, and a sitemap line. Nothing else. So fifteen new subdomains went live with the identical defect, a documented diagnosis sitting on my own website, and no mechanism connecting the two.

Those fifteen then took 682,100 requests and 3,783 MB from the three enumerating crawlers in six days.

I diagnosed it, published it, and shipped it fifteen more times. Writing the article was not the fix.

And This Afternoon the Build Took Atlanta Back

Here is the part I found last, and it is the reason this is going out tonight rather than next week.

I checked all thirty BellyUp category subdomains while writing this. Not one of them currently carries the query-string rule. Every robots.txt is the bare permissive version, 72 to 86 bytes, Allow: / and a sitemap line.

Atlanta's was rewritten today at 21:08 UTC. Charlotte and Tampa at 21:09. Jacksonville, Orlando and St Augustine at 04:08, by the nightly run.

The generator does this unconditionally. It assembles the sitemap, then writes robots.txt from a fixed string, every time it runs. Thirty of the hundred and four sitemap generators in my portfolio write robots.txt the same way. Any rule hand-added to one of those files is live until the next build touches that site, and nothing warns you when it goes.

So the protection that demonstrably took GPTBot to zero for three weeks is gone, on every site that had it, as of this afternoon. If the filter grid is the magnet I think it is, Atlanta's next appointment will show up in the logs within days and I will have an unplanned natural experiment to go with the planned one.

The Receipts

Atlanta network, before the line (2026-08-28 to 2026-09-11) against after (2026-09-12 to 2026-10-02)

  • GPTBot filter-URL requests: 63,649 before, 0 after
  • OAI-SearchBot: 19 before, 0 after
  • Amazonbot: 545 before, 45 after
  • meta-externalagent: 937 before, 10,566 after
  • robots.txt carrying Disallow: /*? recovered from weekly backups stamped 2026-09-11 23:30, 2026-09-19 22:46 and 2026-09-26 04:07

Jacksonville, Orlando and St Augustine, unprotected, 2026-09-27 to 2026-10-02

  • 215 URLs declared across the fifteen category subdomains, counted from their live sitemaps
  • 682,100 requests from the three enumerating crawlers, and 3,783 MB of HTML served
  • 206,356 filter requests against 24 real-page requests on bars.bellyupjax.com
  • 82,913 distinct filter URLs reachable on that one site, against 17 real ones
  • 7,183 to 7,198 requests per hour for eleven consecutive hours, on two separate days
  • ClaudeBot, OAI-SearchBot and PerplexityBot: zero filter requests between them
  • 572,212 of 573,723 GPTBot requests verified inside OpenAI's published address ranges

State as of 2026-10-02

  • 0 of 30 BellyUp category subdomains still carry the query-string rule
  • 30 of 104 portfolio sitemap generators rewrite robots.txt on every run

The Repair Goes in the Generator, and It Is Not a Block

Two things have to change, and neither of them is another hand edit.

The first is where the fix lives. A crawl rule belongs in the generator that writes every site, not in a file the generator overwrites. That is true whichever rule you pick.

The second is which rule to pick, and I am changing my answer. Disallow: /*? worked beautifully against OpenAI and did nothing against Meta, because it asks for cooperation and only some vendors cooperate. Google's own faceted navigation guidance offers an option that does not ask: put the filter state after a # instead of a ?.

A fragment is never sent to the server. bars.bellyupjax.com/#discovery=late-night and bars.bellyupjax.com/ are the same request as far as the server, and therefore every crawler, is concerned. The filter still works for a visitor, because the browser reads the fragment and the page filters itself. There is nothing to canonicalise, because there is only ever one URL. Nothing is blocked, nothing is hidden, and nothing returns anything other than 200. A crawler that ignores robots.txt has no grid to find.

The same guidance adds a repair I need regardless. A filter combination matching no venues should return HTTP 404, not a 200 with an empty-state message. I tried three nonsense combinations on the Jacksonville site while writing this and all three returned 200 and a full page.

So the work is two changes in one generator rather than thirty hand edits, which is also the only version of this that survives the next build.

The Crawler Reads Your Structure, and Your Build Rewrites Your Structure

I went into this week expecting to write that AI crawlers are indiscriminate. They are not. OpenAI read one line and stopped completely for three weeks. That is a well-behaved crawler doing exactly what it was asked.

I came out of it with two less comfortable findings instead.

Compliance is a vendor property, not a crawler property. The same line that bought total protection from OpenAI bought none from Meta, so any plan that depends on robots.txt is only as good as the list of companies that implement it, and that list is not published anywhere I can read.

And a fix is only as durable as the file it lives in. Mine lived in a generated file for three weeks and then quietly stopped existing, on thirty sites, on an ordinary Friday afternoon, with nothing failing and nothing warning me. I only found it because I went looking for the backup to get a quotation right.

The next move is a test, not another article. One subdomain moves its filters behind a fragment in the generator, two comparable subdomains stay exactly as they are, and all three get watched for fourteen days. If the treated site's filter requests collapse while its real pages keep getting read, it goes into every generator that builds a directory.

That experiment is open in the research ledger. There is debate around ambiguity. There is no debating receipts.

Author

Sr. SEO Strategist & Founder

Krisada Eaton

Krisada Eaton is a 25-year SEO Specialist and founder of RealSEOLife.com. He has worked across Fortune 500 companies and independent businesses, with a current focus on AI-ready architecture, digital asset development, and Search Everywhere Optimization.

Content Lab

See Where These Ideas Get Tested

The case studies and experiments are where the ideas in these articles get tested on live sites.