Experiment
Can an Autonomous SEO Research Routine Be Trusted to Say I Don't Know?
A live experiment measuring an autonomous research and publishing routine by what it refuses to promote, invent, or pad when the evidence is not there.
A research routine with a persistent evidence ledger, explicit evidence levels, mandatory primary-source checks, and permission to publish nothing can maintain useful editorial restraint across repeated live runs. Trust will show up as correct refusals, careful promotions, preserved history, and quiet-week discipline, not raw publishing volume.
What is being tracked.
How the test is structured.
RealSEOLife.com's weekly routine reads a private protocol and a current warehouse extract, researches changes in AI and search discovery, updates a public evidence ledger, writes a dated research brief, and may publish one seed article.
Every claim lives at one of four levels: established, external-finding, hypothesis, or demonstrated. Moving a claim up requires new evidence that actually tests it. Repetition and vendor corroboration are not enough by themselves. External claims must be checked against a source page fetched during that run. A quiet week is allowed to end with a short brief and no article.
The scorecard tracks unsupported claims published, bad dates caught, proposed promotions refused or accepted correctly, old entries preserved, duplicate or conflicting entries, quiet-week padding, and whether a property with thin grounding data causes the routine to say it cannot form a useful hypothesis. Operational interruptions are recorded separately from editorial judgment.
What has happened so far.
Run 1 ... September 8, 2026.
The routine published one research brief, one 2,000-word article, and an updated ledger. The ledger grew from four entries to nine without losing any of the original four. Zero claims were promoted.
The stronger receipt was restraint:
- Five proposed outputs or claim moves were declined.
- Three plausible dates from search summaries were rejected after primary-source checks.
- One supported framing stayed at hypothesis because corroboration was not a test.
- One second hypothesis was refused because the warehouse lacked the robots.txt field needed to ground it.
- One property direction was refused because the property was the data point, not the audience.
- One experiment record was refused because its intervention arm was undefined.
- One real development was excluded because it fell outside the brief's discovery scope.
There was also an operational qualification. The run paused for roughly 25 minutes on a publish-permission prompt before Krisada approved it. The research and editorial decision process ran without a review pass. The full production chain was not literally unattended from start to finish.
What would support or challenge it.
Status: active. One good run proves the guardrails can work once. It does not prove they will hold under pressure.
The routine still has three real exams ahead of it. It has never promoted a claim, never faced a week with nothing worth publishing, and never managed a ledger large enough for collisions and ninety-day re-review to matter.
The experiment strengthens if it makes a defensible promotion, publishes no filler during a quiet week, preserves evidence levels past thirty entries, and refuses to invent direction for a property with thin warehouse history. It fails the moment output volume wins an argument against the receipts.
Keep Following the Tests
Move from this open thread back into the full experiment library.