Experiment

Do Crawlers Follow JSON-LD to Resources They Cannot See on the Page?

A controlled follow-up to an AS400Software crawler sequence that requested a catalog and ten JSON datasets in 92 seconds. Those resources appeared on the homepage only inside JSON-LD. The receipt supports discovery through structured data. This experiment is designed to find out whether that is what actually happened.

Krisada Eaton Active ... baseline receipt recorded 2026-09-10 8 views
Hypothesis

When a machine-readable resource is declared inside a page's JSON-LD but has no visible HTML link, a discovery crawler is more likely to request it than a comparable resource that is live but not declared. The observable event is the fetch. Parsing, indexing, citation, training use, and ranking remain separate questions.

Measurement Frame

What is being tracked.

Status Documented
Started September 11, 2026
Duration Active ... baseline receipt recorded 2026-09-10
Branches 3
Setup

How the test is structured.

The originating receipt came from AS400Software.com. On August 10, 2026, a GPTBot-classified client requested the site's catalog and ten JSON datasets in 92 seconds with the homepage recorded as the referer. Current source inspection found all eleven requested URLs inside the homepage JSON-LD and nowhere in the visible homepage links. All eleven resources still return HTTP 200 and valid JSON, covering 369 records.

That sequence is unusually suggestive. It is not controlled. The catalog may have introduced the ten datasets after the first request, cached discovery may have existed elsewhere, and a referer does not prove which page element the client read.

The controlled version will use matched machine-readable resources on the same property. Each pair will have the same content type, response status, robots treatment, sitemap treatment, and publication time. One resource will be declared only in JSON-LD. The other will be live but undeclared. Declaration will be rotated between pairs after the first observation window. Server logs will record verified crawler identity, request time, referer, response code, and request sequence. No ranking or citation claim will be inferred from the fetch.

Observed Signals

What has happened so far.

Baseline receipt recorded September 10, 2026.

  • 49 successful endpoint requests reached 12 AS400Software JSON or AI resources from four crawler fingerprints during the retained observation period.
  • One GPTBot-classified client requested the catalog and 10 JSON datasets in a 92-second sequence on August 10.
  • The homepage was the recorded referer.
  • All 11 requested resources were declared inside homepage JSON-LD.
  • All 11 return HTTP 200 and valid JSON.
  • The dataset layer covers 369 records.

What this supports: structured declarations can sit immediately upstream of a real crawler fetch sequence.

What this does not prove: that the client parsed JSON-LD, understood the records, indexed them, cited them, trained on them, or changed any ranking because of them.

The controlled intervention has not run yet.

Decision Thread

What would support or challenge it.

Status: active. The receipt is strong enough to justify the experiment and not strong enough to skip it.

If declared resources are fetched more often, and the advantage follows when declaration rotates between matched pairs, JSON-LD discovery becomes a demonstrated behavior in this environment. If the undeclared controls are fetched at the same rate, the August sequence had another discovery source.

The measurement stops at the request. Anything beyond that needs its own receipt.

Live Test Property View Property
Experiment Lab

Keep Following the Tests

Move from this open thread back into the full experiment library.