SEO Strategy

The Bots That Ask Permission and the Bots That Don't

One vendor, one portfolio, one 28-day window, three bots, three directions. Reporting AI crawler requests as a single number hides the only part that matters.

Krisada Eaton 10 min read 13 views

The AI bot hitting your server hardest is almost never the one deciding whether you get cited. Those are two different jobs, run by two different classes of software, and over the last few weeks Google and Cloudflare have both started documenting them that way.

I did not get there from a blog post. I got there staring at a bot table, noticing that one vendor's three agents had gone three different directions in the same 28 days across the same servers.

ClaudeBot was flat, 41,863 requests in the prior 28 days down to 40,249 in the recent 28. Claude-SearchBot fell off a cliff, 18,586 down to 5,300, about 71 percent. Claude-User went from 921 to 10,924, close to twelve times.

Same vendor. Same roughly 130 properties in the Digital Karma log ingestion. Same window, the 28 days ending 2026-09-08.

Three directions.

Report AI crawler requests as one line on a dashboard and that entire story shows up as a rounding error.

I Was Reading One Number Where There Were Three

My working assumption, for longer than I would like to admit, was that AI crawler traffic was a volume problem. More bots reading you is better than fewer. Watch the line go up.

That is technically true. It also misses the point.

There are at least three separate jobs under the label AI crawler, and they answer to different rules. Collecting content to train or fine-tune a model. Building an index so a search product can answer a question later. Fetching one specific page right now because a person just asked something.

Those are not stages of one process. They are different products with different appetites, and a property can be thoroughly harvested by the first while staying invisible to the third. If all three tracked the same site signals, Anthropic's three agents would have moved together. Flat, collapsing and multiplying is not moving together.

Then the Documentation Caught Up

What made me stop treating this as a curiosity is that Google has now separated the two classes in its own documentation, and filed them in different places.

Google's crawler and fetcher reference lives in its own Crawling infrastructure section at developers.google.com/crawling now, alongside Search Central rather than inside it, framed as shared across Search, Gemini, Shopping, AdSense, News and Gemini Notebook.

Inside it, the common crawlers page lists the familiar set: Googlebot and its image, video and news variants, Storebot-Google, Google-InspectionTool, GoogleOther, Google-CloudVertexBot, Google-Extended.

Google-Agent is not on that page. It is on the user-triggered fetchers page, described as "used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request." That page carries a last updated stamp of 2026-08-19, and this sentence about the whole class: "Because the fetch was requested by a user, these fetchers generally ignore robots.txt rules."

The same page shows Google-NotebookLM as a former agent, "supported until August 2026," replaced by Google-GeminiNotebook, which "requests individual URLs that Gemini Notebook users have provided as sources for their projects."

OpenAI has said a version of this for a while. GPTBot for training, OAI-SearchBot for ChatGPT search, ChatGPT-User for user actions with the note that "Because these actions are initiated by a user, robots.txt rules may not apply," and now OAI-AdsBot for validating pages submitted as ads.

So robots.txt governs one class of machine reader and not the other. That is not a loophole anyone is hiding. It is written down in plain language on pages you can read this afternoon.

Cloudflare Turns the Split Into a Switch on September 15

Cloudflare's developer changelog entry dated 2026-07-01 does the thing the rest of us have been doing badly in spreadsheets. It makes purpose a control surface.

Three independent categories, in Cloudflare's own words. Search is "crawlers that index your content so they can answer questions about it later." Agent is "automated activity acting in real time on a person's behalf, such as chat fetch bots." Training is "crawlers that take your content to train or fine-tune a model." For each one you can block everywhere, block only on pages that display ads, or not block.

Then the date. From September 15, 2026, new domains onboarding to Cloudflare get defaults where "Bots classified as Training or as Agent are blocked on pages that display ads, while Search remains allowed." And the sentence worth reading twice: "Multi-purpose crawlers that combine Search and Training will be affected by the new defaults to block Training."

If your crawler does two jobs behind one user agent, a policy aimed at the unpopular job now catches the popular one too. The pressure runs toward vendors splitting their fingerprints into single-purpose agents, which is exactly what Anthropic's three-way and OpenAI's four-way splits already look like.

Which means the bot table is about to get more useful, not less.

The Fastest Growth Is Happening in the Newest Access Modes

Here is where I have to be careful not to tell a cleaner story than the numbers support.

Across the log ingestion, recent 28 days against the prior 28, the established broad crawlers grew without step-changing. Amazonbot 48,782 to 74,311 across 127 properties. GPTBot 21,554 to 29,202 across 127. OAI-SearchBot 9,403 to 14,816 across 129. Near-universal footprints, growing 35 to 57 percent.

The multipliers came from somewhere else. Google-Agent 1,358 to 9,473, about seven times, across only 5 properties. Applebot-Extended 446 to 5,746, close to thirteen times, across only 4. cohere-ai 1,444 to 12,811 across 16.

So the tidy version of this article is that narrow bots grow fastest. That version is wrong. Claude-User multiplied almost twelve times across 74 properties and GrokBot went 1,121 to 11,590 across 71. Those are wide.

The honest reading is smaller and more useful. Growth is concentrated in the newer access modes rather than in classic bulk crawling, and those modes vary enormously in how many sites they are aimed at. Some sweep. Some get pointed at a handful of places by something we cannot see from here.

And AgentTrust went the other way entirely, 17,268 down to 2,067. Bots retire, get renamed, or get pulled. A number collapsing is not automatically a verdict on your site.

Our Own Portfolio Disagrees With Itself

The uncomfortable detail is in our own logs.

Over the same 28 days, AIWebsiteSystems.com logged 400 Google-Agent requests. RealSEOLife.com, the site you are reading, logged none. Not a small number. No row at all.

Meanwhile RealSEOLife.com logged 1,701 Applebot-Extended requests, its second heaviest bot behind meta-externalagent at 4,374. Applebot-Extended touched 4 properties in total, so this site sits inside Apple's very small set and outside Google's very small set, on the same infrastructure, run by the same person, with the same publishing habits.

I do not know why. That is the honest answer, and I would rather write it than invent a mechanism.

The question is at least well shaped now. Something decides which handful of properties an agentic fetcher gets aimed at, and it is not the same thing that earns a visit from Amazonbot, which apparently visits everybody.

Meanwhile, Retrieval Is Still Selecting on Shape

A paper from this week lands on the other end of the same problem. Counter-GEO-Bench, arXiv 2609.02316, submitted September 2, 2026 by Bing Zheng, Zongyao Zhao and Wenming Yang. They paired 247 human-verified queries with information-preserving and information-distorting rewrites of the same material, then measured how well existing guardrails stop the distorting version from being retrieved and synthesized into an answer.

Granite Guardian, Llama Guard 3 and NeMo Self-Check Fact-Checking reduced attack success rate by at most 5.7 percent relative, and the paper notes Granite Guardian's reduction was not statistically significant. Their own lightweight baseline, C-GEO Guard, managed 47.6 percent.

The line that explains it: "Safety-taxonomy guardrails target policy violations, while GEO misinformation passes through them as fluent informational content."

Read that as a description of the machinery, not an invitation. At the retrieval step these systems are largely selecting on shape. Well-formed, confident material gets picked up because it looks like the thing that answers the question, and clean prose that happens to be wrong is not a policy violation.

One benchmark, one team, three victim models, published days ago and reproduced by nobody. Hold it at arm's length. But it fits what Google keeps saying in its own generative AI guide, last updated 2026-07-10, which spends its energy on being indexed, being crawlable and being eligible for a snippet, and tells people outright to ignore llms.txt files and content chunking.

There is no separate AI ranking layer to work on. There is a retrieval layer that is easy to feed and, for now, hard to filter.

What This Changes About How We Measure

We already made the sibling version of this mistake, in the other direction, and it belongs next to this one.

When eight articles shipped here between 2026-08-16 and 2026-08-21, weekly GSC impressions went from a run of 144, 273, 75, 40, 47 to 1,936, then 2,561, then 2,296. About a fortyfold move. Average position over the same weeks sat at 57.0, 54.8, 55.4, which is to say it did not move. Clicks stayed between 0 and 2 throughout.

That was an eligibility and surface effect. More pages became eligible for more queries at roughly the same depth. Calling it a ranking improvement would have been a lie told with real data, which is the expensive kind.

The crawler side needs the same discipline. One AI crawler number is three measurements wearing a trench coat. Training coverage tells you whether your material is entering corpora. Search coverage tells you whether an AI search product can find you when someone asks. Retrieval coverage tells you whether a live agent acting for a real person came and opened your page.

Different questions, different controls, and only the third one has a human attached.

What we are actually changing

Split the crawler rollups by declared purpose class, not only by bot name, so training, search and agentic fetches are three columns instead of one.

Carry sites_touched next to request volume. A bot at 9,473 requests across 5 properties and a bot at 9,473 across 129 are telling you opposite things.

Add a per-property robots.txt directive extract and join it to those rollups. We cannot test whether permissiveness correlates with anything, because the field does not exist yet.

Lock a clean four-week baseline before 2026-09-15, so the four weeks after Cloudflare's change are comparable to something.

Keep publishing. Nothing here argues for removing anything, and the eight-article burst is still the clearest evidence we have that surface area buys eligibility.

The Question Is Which Door They Came Through

For about a decade the interesting question was whether the bots came at all. Crawl budget, indexation, log files, the familiar set. For anyone doing the work, that one is mostly settled. They come.

The question now is which door they came through, and on whose behalf.

A training crawler reading your page is a bet on your material being useful someday. A search crawler reading it is a bet on it answering a question later. A user-triggered fetcher reading it means a person asked something in the last few seconds and a machine decided your page was worth opening to answer them.

Only one of those has a human attached. It is also the one robots.txt does not govern, the one Cloudflare starts blocking by default on ad-bearing pages for new domains next week, and the one that showed up on 5 of our properties and not on this one.

We do not architect claims. We architect the receipts. This week the receipt is a bot table that refuses to be summarised into one number, and a warehouse column we now know we are missing.

Author

Sr. SEO Strategist & Founder

Krisada Eaton

Krisada Eaton is a 25-year SEO Specialist and founder of RealSEOLife.com. He has worked across Fortune 500 companies and independent businesses, with a current focus on AI-ready architecture, digital asset development, and Search Everywhere Optimization.

Content Lab

Explore Related Research

Browse our documented case studies, experiments, and systems.