Warehouse-First Content: Check the Data Before You Write a Word
A campaign idea is a guess until you check it against your own data. This one got checked first.
A content idea showed up the normal way. A conversation with Kodi, ChatGPT's content-strategy role in this portfolio, proposed a Digital Asset explainer series and a Digital Asset Observatory.
Good idea. Timely topic. Legislation is bringing digital assets into the news cycle, and the portfolio already had sites built to own that ground.
The normal next step is to start writing.
The actual next step was a query against the Digital Karma Data Warehouse.
Not to validate the idea after the fact. To find out, before any content existed, whether it was worth the day it would take to build.
What the Query Actually Asked
Two questions, both answerable from data the portfolio already collects.
First: is anyone already ranking for this topic? Ninety days of GSC data across the proposed sites, plus a wider scan of every 'digital asset' family query across the full portfolio.
Second: is anything already reading these sites on this subject? Ninety days of AI-crawler logs, joined against the same query window.
The first answer was close to zero. Nine thousand nine hundred impressions, twenty-eight clicks, average position in the 40s to 70s across the proposed cluster. Nobody owned this topic yet.
The second answer was the opposite. Thirteen thousand seven hundred eighty-one AI-crawler requests over the same ninety days. ClaudeBot and GPTBot were already reading these sites daily.
The sites were not invisible. They were unread by the instrument that was being used to judge them.
The Plan Changed Once the Data Was In Front of It
The original proposal had two candidates for hosting a Digital Asset Observatory: one investment-framework site, one AI-research site.
The warehouse settled it. The investment-framework site already had the only glossary, directory, and category scaffolding across the whole cluster. A related property had already claimed the measurement-layer role in a post from three days earlier.
The architecture assigned itself once the data was in front of it. That almost never happens when a plan starts from a topic instead of a query.
Nine Sites, Three Different Architectures
Seven sites were in the original plan. Three of them turned out to have no blog system at all, they were marketplace or product templates, not article-based sites.
One needed a single line added to a router allowlist. One already had a generic page template built for exactly this. One got a new page cloned from an existing static-page pattern already live on the site.
A ninth site got added mid-build after it was flagged as growing on its own. The warehouse confirmed real, climbing GSC impressions, and turned up a second finding along the way: that site's report template had been shipping placeholder text for every report it had ever published. Fixed as part of adding its first real one.
Guessing at a single content format across nine sites would have broken at least three of them. Reading the router first cost a few extra minutes per site and avoided that entirely.
Before You Write the Next Campaign
Check the data first
- Pull 90 days of GSC data for the exact query cluster you are targeting
- Pull AI-crawler logs for the same sites and window
- Compare the two. A gap between them is the actual opportunity
Let the data assign roles
- Check which property already has the structural scaffolding for a hub page
- Check for existing claims on the same positioning elsewhere in the portfolio
Read before you write
- Confirm each site's actual content architecture before assuming a template
- Validate every file locally before it goes live
- Verify every URL live after publishing, not just that the upload succeeded
The Idea Was Good. The Query Made It a Plan.
Kodi's proposal was sound content strategy on its own. It became something stronger once it was checked against data the portfolio already had sitting in a warehouse.
Thirteen pieces of content shipped across nine sites in one day, each one aimed at a specific, verified gap instead of a general hunch about a trending topic.
The campaign is not proven yet. Search visibility takes weeks to move, and that is being tracked as a live experiment, not assumed as a result.
What is already proven is cheaper than the campaign itself: the data was there before the content was, and checking it first is the difference between publishing on a guess and publishing on a gap you can actually see.
Explore Related Research
Browse our documented case studies, experiments, and systems.