Experiment

The Block Came Off by Accident. Does the Crawler Come Back?

An unplanned before-and-after. A robots.txt rule that had held GPTBot to zero filter-URL requests for twenty-one days was removed from thirty BellyUp category subdomains by the nightly sitemap build on 2026-10-02. The before state is three weeks of measured compliance. The question is whether the filter grid pulls the same crawlers straight back in.

Krisada Eaton Ongoing ... first checkpoint planned for 2026-10-09 7 views
Hypothesis

If a combinatorial filter grid is what draws bulk training crawlers rather than anything about the site's content, then removing the query-string block will bring the enumeration back on its own, with no other change to the sites. GPTBot should resume filter-URL requests within days rather than weeks, in the same single-day rate-limited session shape measured on the Jacksonville network, and the volume should scale with each subdomain's reachable combination count rather than with its real page count. A null result, meaning the crawlers do not return to the grid within four weeks despite the block being gone, would mean something other than reachability governs enumeration and would weaken the magnet reading considerably.

The Scorecard

What this test is watching.

A line in a small text file told AI bots to skip the filter links on these websites. It worked for three weeks. Then the site's own build routine erased the line by mistake. We are watching to see how quickly the bots start crawling tens of thousands of filter pages again.

Status Ongoing
Started October 2, 2026
Duration Ongoing ... first checkpoint planned for 2026-10-09
Branches 4
Setup

How the test is structured.

This experiment was not designed. It was created by a build.

On 2026-09-11 at 23:30 UTC a one-line rule, Disallow: /*?, was added to robots.txt on the BellyUp Atlanta category subdomains after a crawl trap was diagnosed and published on 2026-09-09. The exact file was recovered from weekly backup archives stamped 2026-09-11 23:30, 2026-09-19 22:46 and 2026-09-26 04:07, so both the wording and the dates are evidence rather than recollection.

The before state, measured across the full raw log retention window up to the change (2026-08-28 to 2026-09-11) against the period after it (2026-09-12 to 2026-10-02), is a clean per-vendor split on the Atlanta network:

  • GPTBot: 63,649 filter-URL requests before, 0 after
  • OAI-SearchBot: 19 before, 0 after
  • Amazonbot: 545 before, 45 after
  • meta-externalagent: 937 before, 10,566 after

On 2026-10-02 the per-site deploy/generate-sitemap.php regenerated robots.txt across all thirty BellyUp category subdomains, writing a fixed string containing User-agent: *, Allow: / and a sitemap line with no merge and no conditional. The Atlanta files were rewritten at 21:08 UTC, Charlotte and Tampa at 21:09, and Jacksonville, Orlando and St Augustine at 04:08 by the nightly cron. Checked the same day, none of the thirty carries the rule any longer. Thirty of the portfolio's one hundred and four sitemap generators of that name write robots.txt the same way.

Nothing else changed. Canonical tags on every filter URL still point at the unfiltered hub and have done throughout. No page was added, removed, blocked or hidden. The filter grids themselves are unchanged.

The comparison arm is already in hand. Over 2026-09-27 to 2026-10-02, the fifteen Jacksonville, Orlando and St Augustine subdomains, which never received the rule, absorbed 682,100 requests and 3,783 MB from the three enumerating crawlers, including 206,380 requests to bars.bellyupjax.com against 22 declared sitemap URLs. That is what an unprotected grid of this kind attracts on this infrastructure.

Measurement is per subdomain per day: filter-URL requests split by crawler, clean-URL requests over the same days, bytes served, and the reachable combination count. The Atlanta subdomains are the treated group in the sense that they lost a protection. Jacksonville, Orlando and St Augustine are the already-unprotected reference. First checkpoint is planned for 2026-10-09, while the raw request evidence sits well inside the 35-day retention window.

Observed Signals

What has happened so far.

Baseline recorded 2026-10-02, the same day the block was removed. No post-removal observation yet.

Atlanta network before-state, carried forward from the compliance measurement above:

  • 63,649 GPTBot filter-URL requests in the fifteen days before the rule
  • 0 GPTBot filter-URL requests in the twenty-one days with the rule in place
  • 575 GPTBot clean-URL requests during those twenty-one days, so the crawler kept visiting real pages while skipping the grid entirely
  • 10,566 meta-externalagent filter-URL requests during the same twenty-one days, since that crawler never honoured the rule

One question this experiment does not yet answer, and should not be read as answering: whether the block redistributed crawl attention toward real pages. Clean-URL requests per day on the Atlanta network did rise after it went in, GPTBot from 13.3 to 27.4, OAI-SearchBot from 12.3 to 22.2 and Amazonbot from 4.5 to 8.1. Portfolio-wide GPTBot volume rose roughly twentyfold across the same window, which is far more than those increases, and the two candidate control sites, Charlotte and Tampa, were not being crawled at all before the change and therefore supply no before-window. With a rising baseline and no valid control, that rise cannot be attributed to the block. It stays an open question.

The removal does supply the control that was missing. If the grid is the magnet, Atlanta's filter requests should return while its real pages continue roughly as they are, which separates the two effects on the same site with the same crawlers.

Decision Thread

What would support or challenge it.

Status: active, baseline only.

This is the cheapest useful experiment available right now because the intervention has already happened and the before-state is three weeks of measured compliance rather than an estimate.

If GPTBot resumes filter-URL requests on Atlanta within days, in the single-day rate-limited session shape already measured on Jacksonville, then reachability governs enumeration and the magnet reading holds. The practical consequence is that the repair has to move into the generator and should not depend on a text file the generator rewrites.

If the crawlers do not return to the grid within four weeks despite the block being gone, then something other than reachability is selecting what gets enumerated, and both the magnet reading and the planned fragment experiment need rethinking before either gets built into a generator.

Either way the durability finding stands on its own and does not need this result. A crawl directive stored in a generated file has an undeclared expiry, nothing warns when it lapses, and no current warehouse table records what any site was serving on a given date.

Live Test Property View Property
Experiment Lab

Keep Following the Tests

Move from this open thread back into the full experiment library.