Case Study

AI Search Visibility Using SEO Datasets

Thirty-five days of server logs across 190 websites show AI crawlers and search engines fetching published SEO datasets within days, following page links to the files, and more than doubling their catalog fetches after Dataset structured data went on every page. Clicks did not move, and no log can show a citation.

Krisada Eaton 13 views
Proof Point

Crawlers took the file the page pointed to.

Case Receipt

The claim, evidence, limits, and open loop.

Retrieval confirmed ... citation open
Starting Condition

What was true before the work

Before September 2026, AI Symantix published its measurements as prose. The portfolio declared datasets only inside each site's catalog file, and no one had checked whether crawlers took the data files a page linked, or whether declaring datasets on every page changed anything.

Intervention

What changed

AI Symantix published five SEO datasets between September 7 and 12, each as a page, a JSON file and Dataset structured data naming the file. On September 23 every rendered page on all 190 portfolio sites began carrying DataCatalog and Dataset JSON-LD pointing at its catalog, under Digital Karma Federation v8.2.

Result

What the evidence showed

Every AI Symantix data file was fetched within days, and the referers show Googlebot, GPTBot and PetalBot following the page link to the file. Across 190 sites the dataset layer drew about 21,700 requests against 7.09 million content requests. OAI-SearchBot spent 3.9 percent of its visits on the dataset layer against 0.04 percent for GPTBot. After September 23 the number of sites where OAI-SearchBot fetched the catalog went from 11 to 43, GPTBot from 54 to 87 and Meta from 3 to 19.

Search Console

What Google reported

Google showed seven AI Symantix dataset pages between August 10 and October 8, 2026, for 226 impressions and zero clicks. The CTR dataset page sits at position 58 and its top query is the pasted header row of a Search Console export. Site-wide weighted position moved from 88.2 in mid June to 30.4 in early October while weekly clicks stayed between zero and two.

Server Logs

What the server logs reported

Thirty-five retained days, September 6 to October 10, 2026. Googlebot fetched /data/lab/exposure-velocity.json 18 times with /ai-symantix-lab/ as referer. GPTBot fetched the search conversion baseline on AI Symantix and DatasetSEO on September 12 with signalarchitectgroup.com as referer. Bots that could not be named made 11,969 dataset-layer requests and reached the catalog on 176 sites.

Limits

What this does not prove

A successful request proves a request reached the server, not indexing, citation, training use or a reader. The before and after windows differ by one day and crawler schedules are not observable, so the catalog jump is a sequence rather than a proven cause. Human counts may include our own visits. Raw AI crawler volume in October is dominated by BellyUp facet crawl traps and was excluded from every finding.

Current Status

Where the case stands now

Retrieval of published SEO datasets is confirmed across search engines and AI crawlers. The wider catalog reach after September 23 is logged as a strong hint. No citation has been observed, and no finding here claims one.

Next Checkpoint

What gets measured next

Re-count distinct catalog sites per AI crawler for October 11 through November 10 to see whether the post-September 23 reach held, read the November 1 Search Visibility Index for AI Symantix, and start a fixed-question citation test for the five datasets in at least two AI tools.

Measurement Frame

What was measured.

Three readings from the same 35 days of logs. Each one is a retrieval fact, and none of them is a citation.

The link is the route
18 of 18 fetches
Googlebot requests for one JSON file, all with the Lab page as referer

The crawler did not find the file by guessing. It read the page, then took the file the page linked. GPTBot and PetalBot did the same on other files.

Search crawlers want the catalog
3.9% versus 0.04%
share of visits on the dataset layer, OAI-SearchBot versus GPTBot

The crawler OpenAI uses to answer questions spends about one visit in 25 on data files. The one it uses for training spends about one in 2,300.

Declaring on every page widened reach
11 sites to 43 sites
OAI-SearchBot catalog fetches, 17 days before and 18 days after September 23

The same jump appears for GPTBot and Meta. It is a sequence with a date, not proof of cause, and the next monthly reading will say whether it held.

The dataset layer was three tenths of one percent of 7.09 million content requests. Small, fast, and measurable is still the honest description.

Summary

The short version.

AI Symantix published five SEO datasets in September 2026 and the whole portfolio began declaring datasets on every page on September 23. This case study joins the warehouse server logs and Search Console data to show who fetched the files, how they found them, what changed after the portfolio-wide declaration, and what Google did with the dataset pages.

Plain English

Explain it to me like I'm ten.

We put our own SEO numbers online as public data files. Search engines and AI crawlers came and took the files within days, and the logs show they followed the link on the page to get there. After we labeled the datasets on every page, AI crawlers picked up the catalog on more than twice as many sites. None of that turned into Google clicks yet, and the logs cannot tell us whether any AI quoted us.

Sites Where Each AI Crawler Fetched the Catalog

Distinct portfolio sites with a successful /ai/catalog.json request, BellyUp sites excluded. Before is September 6 to 22; after is September 23 to October 10, 2026.

Three companies' crawlers reached the catalog on more sites after every page began declaring its datasets. Only ClaudeBot stayed flat.

GPTBot before 54
105 fetches
GPTBot after 87
127 fetches
OAI-SearchBot before 11
14 fetches
OAI-SearchBot after 43
44 fetches
meta-externalagent before 3
10 fetches
meta-externalagent after 19
28 fetches
ClaudeBot before 5
5 fetches
ClaudeBot after 6
7 fetches

Sites with a catalog fetch

Context

What was the system?

In September 2026 AI Symantix published five SEO datasets in a row: a click-through-rate curve by Google position, a monthly Search Visibility Index, two AI bot visibility studies and a search-to-inquiry baseline. Each one shipped three ways at once, as a readable page, a downloadable JSON file, and Dataset structured data naming that file. All five were built from the Digital Karma Data Warehouse, which imports Search Console data and classified server logs for every site in the portfolio every night.

Then on September 23 the whole portfolio joined in. Digital Karma Federation v8.2, the standard kept at DigitalKarmaWeb.com, made every rendered page on every site carry a DataCatalog node and Dataset nodes pointing at that site's catalog file. That gave the warehouse a clean before and after inside its 35-day log window.

The question this case study answers is the one everyone asks after publishing data for AI: did anything come and get it, and did it matter? The site-level write-up with every table is AI Search Visibility Using SEO Datasets on AI Symantix. This entry is the portfolio receipt.

Methodology

What changed?

Server logs cover September 6 through October 10, 2026, the 35 days the warehouse retains. Search Console data for AI Symantix runs from June 16 through October 8, 2026. The portfolio means the 190 websites in the Digital Karma registry that served content in the window; domains outside the portfolio in the warehouse sites table were excluded by name.

The dataset layer means successful GET requests to five kinds of file:

  • The catalog at /ai/catalog.json
  • The manifest at /ai/manifest.json
  • The llms.txt and llm.txt guidance files
  • The federation files and any other JSON under /ai/
  • Public JSON data files under /data/, /content/ and /datasets/

Crawlers were named from the warehouse fingerprint table and grouped by the job their owner publishes for them: search, AI search, AI training, or fetching on behalf of a person. Requests classified as a bot with no matching fingerprint are reported separately as bots that could not be named.

Content page totals for the same window come from the warehouse's own portfolio request cache. The before and after comparison counts distinct sites on which each AI crawler fetched /ai/catalog.json in the 17 days before September 23 and the 18 days after, with every BellyUp city site excluded because those finders became facet crawl traps in October. Weighted position is position weighted by impressions. Referers were read straight from the log rows.

Findings

What did the evidence show?

Crawlers took the file the page pointed to. Googlebot fetched the Exposure Velocity JSON on AI Symantix 18 times in 35 days, and every fetch carried the Lab page as its referer. GPTBot read the five-site cohort page and then took the cohort file. PetalBot did the same on six files. The CTR dataset file was fetched the same day it was published, GPTBot read its page three days later, and Bingbot had the file within eight days.

The link travelled between sites too. On September 12 the Search Visibility to Inquiries study was linked from Signal Architect Group. That day GPTBot fetched the baseline file on AI Symantix and the republished copy on DatasetSEO.com, both with signalarchitectgroup.com as the referer. It is the same pattern as the August receipt on AS400Software.com, where a GPTBot client took eleven resources declared only in JSON-LD in 92 seconds.

Across 190 websites the dataset layer drew about 21,700 successful requests against 7.09 million content page requests, three tenths of one percent of the crawl. The biggest share, 11,969 requests, came from bots the warehouse could not name. People made 3,390, including 1,853 reads of llms.txt and llm.txt from 1,469 different addresses. Search engine crawlers made 1,463, AI training crawlers 1,384, AI search crawlers 244, and AI tools fetching for a person 11.

Share of visits is where the pattern shows. OAI-SearchBot, the OpenAI search crawler, spent 3.9 percent of its visits on the dataset layer. GPTBot, the OpenAI training crawler, spent 0.04 percent. Applebot, Bingbot and PetalBot sat between 1.0 and 1.5 percent. Googlebot sat at 0.37 percent. The training crawlers from Meta, Anthropic and Amazon sat at 0.02 to 0.03 percent. Crawlers that answer questions want the catalog. Crawlers that collect text want pages.

Catalog fetches more than doubled after Dataset structured data went on every page. With the BellyUp sites excluded, OAI-SearchBot fetched the catalog on 11 sites in the 17 days before September 23 and 43 sites in the 18 days after. GPTBot went from 54 sites to 87. Meta's crawler went from 3 to 19. Across the whole portfolio GPTBot's busiest catalog day before the change touched 11 sites; on September 24 and 25 it touched 28 and 30.

Google listed every AI Symantix dataset page as an ordinary page and sent zero clicks to any of them. The CTR dataset page has 54 impressions at position 58. Twenty-seven of them came from one query: "top queries,clicks,impressions,ctr,position". That is the header row of a Search Console export pasted into Google. On DatasetsMaker.com the SEO Keyword and Ranking Dataset page has 797 impressions and one click, with "seo datasets" at position 62 and "seo dataset" at position 74.

AI Symantix search visibility moved over the same months, though the datasets cannot be separated from the checker, the glossary cluster and the page consolidation that landed in the same weeks. Weighted position went from 88.2 in the week of June 16 to 30.4 in the first week of October. The AI search visibility query family went from 89.9 to 20.9. Clicks stayed between zero and two a week, which is exactly what the site's own CTR curve predicts for 900 impressions a week at position 30.

The largest number in the warehouse this month was left out on purpose. AI crawlers made 763,949 requests across the portfolio in September and 3,571,310 in the first 11 days of October, and the twelve busiest October sites are all BellyUp finder lenses looping through filter combinations. Raw AI crawler volume appears nowhere as a finding here.

What We Kept

What can be reused?

Publishing SEO datasets changes the retrieved part of AI search visibility and nothing else you can see from a server. The file gets fetched within days, the crawler follows the page link to reach it, and the AI search crawlers show a measurable appetite for the catalog that the training crawlers do not. Put the catalog on every page and the number of sites those crawlers reach it on doubles.

What it does not do is produce clicks or prove a citation. Google lists a dataset page like any other page, at the position the page earns, and the Search Console Dataset report only counts pages that carry the Dataset label. A fetch is the first of three parts. The other two still have to be tested by asking AI tools the question and reading the answer.

The measurement itself is the method. Logs show what Search Console cannot, Search Console shows what logs cannot, and the dataset that AI Symantix published to explain click-through by position ended up explaining its own click famine.

Evidence Library

Keep Browsing the Case Library

Move from this evidence file back into the full proof system.