Field Note
An Unattended Routine Published to a Live Site. The Receipt Is What It Refused to Do.
A weekly research routine ran with no review step in front of it and wrote three files straight to production. The useful record is not the article it produced. It is the five things it declined to produce, and the one guardrail that has still never fired.
The setup.
The weekly research routine on RealSEOLife.com is automated end to end. It reads a private protocol file and a warehouse extract, researches what changed in AI and search discovery, updates a running evidence ledger, and publishes a brief, the updated ledger, and exactly one seed article. Nothing sits between it and the live site. The ledger it maintains holds every claim at one of four levels: established for platform documentation, external-finding for a credible third party we have not reproduced, hypothesis for a working theory stated so it could be shown false, and demonstrated for something measured in our own warehouse. The whole point of those levels is to stop a working theory from drifting into a stated fact because it got repeated three weeks running. The first real run fired on 2026-09-08 against a warehouse extract generated twenty minutes earlier.
The finding.
The run published a brief with four developments, one new hypothesis, and a seed article on the split between crawlers that obey robots.txt and user-triggered fetchers that Google documents as generally ignoring it. All of that landed and verified live. The more instructive half is the refusals. It declined to promote an existing hypothesis that had picked up supporting vendor documentation, on the grounds that corroborating the framing of a claim is not the same as testing the claim. It declined to open a second hypothesis about robots.txt permissiveness because the warehouse extract carries no robots.txt field to test it against, and recorded that gap inside the hypothesis it did open. It declined to write a direction for one property because that property was the week's data point rather than its audience. It declined to publish an experiment record whose observational arm was defined and whose intervention arm was not. It declined to carry a fifth verified development because it was about search market structure rather than website discovery. Separately, three date claims arrived from search result summaries and none survived contact with a primary page: one placed an OpenAI crawler documentation change in September 2026 when the credible dated report puts it at 2025-12-09, one garbled the European Search Dataset Licensing Program timeline that Google's own page states as measures adopted 2026-07-16, licensing agreement available 2026-09-17, and samples available 2026-11-16, and one asserted a September 2026 origin for Google-Extended that Google's crawler documentation does not support. All three would have read as entirely plausible on a live site.
| Level | Before | After | Added this run |
|---|---|---|---|
| established | 0 | 3 | 3 from primary vendor documentation |
| external-finding | 0 | 1 | 1 unreproduced preprint |
| hypothesis | 2 | 3 | 1 new, status open |
| demonstrated | 2 | 2 | 0 |
| Total entries | 4 | 9 | 5 |
| Promotions | 0 | 0 | 0 |
| Declined | Reason |
|---|---|
| Promote crawler-mode-split-hypothesis | New vendor documentation corroborated the framing of the claim, which is not a test of the claim |
| Open a second hypothesis on robots.txt permissiveness | The warehouse extract has no robots.txt field, so it could not be grounded in anything |
| Write a property direction for AIWebsiteSystems.com | The property was the week's data point, not its audience, and no defensible guidance exists yet |
| Publish an experiment record | The observational arm was defined; the intervention arm was not |
| Carry a fifth verified development | Real and primary-sourced, but about search market structure rather than website discovery |
The change.
Nothing was removed and nothing was rewritten. The ledger went from four entries to nine, all four originals verified present after the write, with zero promotions. The routine's own gap is now the next piece of work: the warehouse has no per-property robots.txt extract, so the most obvious confounder on this week's hypothesis cannot be controlled for, and adding that join is the blocker on the experiment rather than a nice-to-have. For the rotating version of this routine being considered for the smaller properties, two things travel with it and one thing does not. The evidence ladder travels, because it lives in the ledger file on the site rather than in the routine's instructions, which means it accumulates across runs instead of resetting whenever the instructions get edited. The rule that every external claim must come from a page fetched during that run travels, and it is what caught all three bad dates. What does not travel is the grounding data. Several of the smaller properties have very thin warehouse history: seo.krisada.com carries AI crawler rows only from the week of 2026-08-24, and signalarchitectgroup.com carries only three weeks of Search Console rows. A routine pointed at a property like that can still research and still publish, but it cannot honestly form a hypothesis, and the correct output is a brief with no hypotheses rather than an invented one.
Three guardrails remain untested and each has a specific trigger. Promotion has never fired, because no entry has moved up a level; the test arrives the week the agentic-fetcher experiment returns real numbers and the routine has to decide whether hypothesis becomes demonstrated. Restraint on a dead week has never been tested, because this week had enough genuine material that publishing nothing was never on the table; the test is the first week with no real development, where the correct behaviour is a short brief that says the week was quiet and no padding. Integrity at scale has never been tested, because nine entries cannot realistically collide; the test is somewhere past thirty entries, with the ninety-day re-review rule actually forcing old claims back under review. Until all three have fired at least once, this counts as a working pipeline with unproven judgement, not a finished system.
More real work, shown
Every field note is one finding from a live website, the data behind it, and the metric we are watching next.