ClaudeBot Changed Its Mind About Filter URLs. Only on the Sites It Just Met.
1,367 filter requests in 32 days. Then 261,124 in four days. Every one on a site ClaudeBot had met that same week.
ClaudeBot started crawling our filter URLs on October 6. It had not requested a single one in the 32 days before that.
Over the next four days it fetched 261,124 filter addresses across the portfolio and pulled 1.8 GB off the server.
Here is the part that matters. It only did this on sites it met that week. On the sites it had already been crawling since September, it kept making the same seventy or eighty clean requests a day and never touched the grid once.
Same vendor. Same bot. Same four days. Two different behaviours, split by nothing except when the bot first found the host.
Last Friday I published a piece here arguing that filter-grid crawling looks like a vendor-level decision. GPTBot did it. Amazonbot did it. Meta's crawler did it. ClaudeBot was the clean counter-example that made the pattern readable at all.
Six days later the counter-example switched sides.
What ClaudeBot Was Doing Before October 6
A filter URL is an address with a question mark in it. On our city guide sites one looks like /?dietary=vegan_options&discovery=brunch&neighborhood=downtown. Three dropdowns, one address each, tens of thousands of combinations.
Three crawlers have been walking those combinations for weeks: GPTBot, Amazonbot and Meta's crawler. ClaudeBot was not one of them.
Our raw logs start on September 4. From that day through October 5, ClaudeBot made 1,367 filter requests across the whole portfolio. Thirty two days. Some days it made none at all. The busiest single day was 235.
Over the same stretch it made between 567 and 4,629 clean requests a day. So it was busy. It just read real pages.
Our raw logs keep 35 days by default. September 4 is as far back as I can look. Before that I only have rollups, and the rollups strip the question mark off the address. That is exactly the detail this whole story turns on.
Four Hours to 8.4 Requests a Second, Then a Ceiling
The ramp is the part I did not expect.
At 19:00 UTC on October 6 ClaudeBot made 167 requests across the portfolio, which is an ordinary hour for it. Then it opened the throttle.
The Ramp, Hour by Hour
A Fixed Rate Is Not an Accident
After the peak, ClaudeBot backed off to about 1,850 requests an hour and stayed there for roughly twenty two hours. It never drifted more than about eighty requests either side of that number.
Half a request a second, held flat for most of a day. That is not a crawler running out of work. That is a crawler obeying its own ceiling.
GPTBot did the same thing last week at a different number. Its ceiling was 7,196 requests an hour. That works out to 1.999 a second, and it held within nine requests an hour across twenty two hours on two separate days.
Two vendors. Two fixed rates. Same shape.
Then on October 8 it ramped a second time. The 12:00 hour hit 19,056 requests before it settled back to the same flat band by evening. By midday on October 9 it was back to background.
Thirty Published Pages. 68,643 Addresses Requested.
Take one site. over50.bellyupstpete.com lists 30 URLs in its sitemap. Thirty real pages.
In four days ClaudeBot requested 68,643 different addresses on it. Those addresses resolved to 36 real pages. Every single request came back 200.
That is 472 MB of transfer to read thirty pages worth of writing.
Savannah's Over 50 guide got 48,805 distinct addresses against 24 published pages. Atlanta's got 29,882 against 27.
Nothing is broken here, which is the uncomfortable part. The canonical tag on every filtered view points at the unfiltered page, the way it is supposed to on all of these sites. The crawler is simply treating a set of dropdowns as a page generator, and the dropdowns are happy to oblige.
How I Know It Was Really ClaudeBot
A user agent is a claim, not an identity. Anybody can type ClaudeBot into a header, and plenty of traffic does.
Anthropic publishes its crawler ranges as a JSON file at claude.com/crawling/bots.json. I fetched it during this run. It carries 38 IPv4 prefixes and a creation time of October 7, 2026.
Those four days carried 271,991 requests wearing the ClaudeBot name. 271,866 of them came from addresses inside that published list. That is 99.95 percent. The remaining 125 came from eight addresses that are not on it, which is the usual borrowed-identity tail.
Almost all the heavy traffic came from four addresses inside 216.73.216.0/22. The ARIN registry lists that block as AWS-ANTHROPIC, registered to Anthropic, PBC.
So the identity holds. This was the real crawler, running at a rate it chose.
Jacksonville Never Got Touched
Now the part that breaks the simple story.
ClaudeBot has been crawling the six Jacksonville guide sites every day since September 18. It was on Jacksonville during all four days it spent walking Savannah, St Pete and Atlanta. It made between 62 and 93 clean requests a day there.
Filter requests to Jacksonville over those four days: zero.
Same grid. Same generator. Same permissive robots.txt. I checked every one of those files this run. They are the same four lines, with Allow: / and no Disallow at all.
The split lines up with exactly one thing. Every host ClaudeBot first crawled on or after October 6 got walked. Every host it already knew did not. Eighteen guide sites in the first group, and all eighteen were walked. Twenty hosts in the second group, and the only filter requests any of them saw were twelve on the Atlanta hub.
Charleston is the gap worth watching. It launched the same day as Savannah and St Pete, and ClaudeBot has touched its six guide sites twice each and gone no further. Meta's crawler walked Charleston hard in the same window, 294,462 filter requests across the six sites. ClaudeBot simply has not arrived yet.
That makes Charleston the next real test, and it is one I do not have to build.
Three Things We Changed That Week, and Why None of Them Fit
I run these sites, so I have to rule myself out before I blame a crawler.
Our changes against the crawler timeline
What we changed on the BellyUp! network
- October 9, 01:00 UTC. A header and Top Picks change shipped to all 54 guide sites.
- October 9, 01:50 UTC. A user-agent block for a fake-browser botnet went into 132 live folders.
- October 9, 15:03 UTC. A navigation order change shipped to the same 54 guide sites.
When ClaudeBot actually started
- October 6, 20:34 UTC. First filter request, on the Savannah bars guide.
The Dates Rule Us Out
All three of our changes landed more than two days after the behaviour started. I read the write times on the live server rather than trusting my own notes.
All three also shipped to every city, Jacksonville included. A change applied everywhere cannot explain a behaviour that happened in three places and not the other six.
The robots files did not change either. They were permissive before, they are permissive now, and they are identical across the cities that got walked and the cities that did not.
So the start date belongs to the crawler.
What This Changes If You Read a Crawler Report
If you keep a note anywhere that says a particular crawler respects your filter rules, put a date on it. Mine was six days old and already wrong.
Three changes to how I read this data
Three changes
- Carry the first-seen date for every bot and host pair, not only the request count. The split in this story is invisible without it.
- Report filter requests and clean requests as two separate numbers, always. One hides the other completely.
- Re-check any claim about a crawler's manners before repeating it, including my own from last week.
Why the Rollups Hid It
There is a measurement trap sitting underneath all of this, and it is worth naming.
Our crawler rollup stores a hash of the address with the question mark stripped off. So 68,643 filter addresses collapse into 36 real pages, and the daily total looks like enormous interest in a handful of pages.
Read only the rollup and this week looks like a popularity spike. Read the raw log and it is a dropdown being counted.
A crawler request proves a request arrived. It does not prove a person read anything, and it does not prove a page was indexed, cited or used. Those need their own numbers.
Good Behaviour Is a Snapshot, Not a Setting
The useful thing about being wrong in public is that the correction gets a date too.
On October 2 I wrote that filter-grid crawling looked like a stable vendor policy, and that ClaudeBot sitting on the clean side of it was the detail that made the pattern readable. My own logs contradicted that four days later, on my own sites.
Part of the claim survives. Whether a crawler walks a dropdown still groups by company rather than by how we have the crawler filed. ClaudeBot and GPTBot are both filed as training crawlers in our own records, and for a month they behaved completely differently on the same addresses. Now they behave alike. The filing never predicted either state.
What died is the word stable.
Crawler behaviour is not a property of a vendor. It behaves like a release. It ships on a date, it has a scope, and the scope can be narrower than every site you own.
So if you are protecting a faceted directory, the fix that holds is the one that does not rely on a crawler's manners. Put the filter state in the piece of the address after a hash mark and the server never receives it. There is nothing to request, nothing to obey, and nothing blocked or hidden. That change is already written up as an open test here and it is still the next thing to run.
In the meantime, 1.8 GB left the server in four days to deliver about thirty pages worth of reading to a bot that had politely ignored them the week before.
I would rather know that than assume it.
See Where These Ideas Get Tested
The case studies and experiments are where the ideas in these articles get tested on live sites.