What you will be able to do
- Set up a ledger repo with one record per SEO experiment
- Split page families into treated and control groups with a seeded random assignment
- Read a diff-in-diff verdict and decide keep, extend or roll back
- Schedule a daily pulse that snapshots Search Console data and wakes due experiments
- Tell which fix types your site can measure at all, and which it cannot
Key takeaways
- An SEO fix counts only when a randomised control group of untouched pages moved less than the treated pages.
- Write the hypothesis, metric, page groups and kill rule into the ledger before the change, and never edit them after.
- A script computes every metric and verdict; Claude Code runs the steps and reads the result.
- Inconclusive is a normal verdict on a small site: the loop extends once to day 59, then closes.
- A daily pulse is the clock: experiments park on a due date and are woken at day 15 and day 31.
What the SEO experiments loop does
SEO experiments only tell you something when the pages you changed moved more than a group of untouched pages. We run our site's SEO fixes through that rule. Claude Code writes the experiment record before it touches a page, changes a randomised treated group, leaves a matched control group alone, and a script compares the two at day 15 and day 31. ConvOps runs it all as scheduled tasks with workflows, and a human merges every change.
A pre-registered SEO experiment is a page change whose hypothesis, metric, page groups and kill rule are committed before the change ships, and never edited after. A page family is one page plus its locale variants (en, es, de). Families are the unit the loop assigns to treated or control.
Every screen and number below uses example data for Tidewell, a fictional invoicing app at tidewell.example.com.
| Part | What it is |
|---|---|
| Trigger | Four schedules. Nobody starts a run by hand |
| Inputs | Search Console page data, AI crawler fetches and human visits from site analytics |
| Agent runtime | Claude Code, one isolated session per run |
| Runner | ConvOps tasks, workflows and schedules |
| Outputs | Snapshots, experiment records, research and strategy memos, all in one ledger git repo |
| Human gates | Every merge to the site, and every strategy memo |

Four schedules drive it: a daily pulse, a weekly planner, and monthly research and strategy.

Why we built it: one page is noise, and agents invent numbers
Most small-site SEO tests lie to you: one page swings by a third with nothing changed. Google updates, seasonality, crawl timing and site-wide dips all move a page on their own. A before and after chart of one page cannot tell a fix from a quiet week.
The answer is a control group from the same site, assigned at random at the same moment. Both groups face the same conditions, so the gap between them is the effect. Google's testing guidance says the time a reliable test needs depends on how much traffic the site gets (Google Search Central, 2025(opens in new tab)). This loop never changes a URL or adds a redirect, and a small site gets 28 measured days plus one extension.

The second problem is the agent. Ask an agent to "check the data" and it computes ratios, compares them with benchmarks it made up, and recommends changes. Ours did exactly that. So a script computes every metric and verdict, and Claude Code runs the steps and reads the results.
If you want a scheduled loop like this, with steps, gates and human approvals, see how ConvOps runs it(opens in new tab).
How it runs: four schedules, five workflows
An experiment loop is only as honest as its clock, so nothing here waits for a person to remember it.
| Schedule | When (UTC) | What it does |
|---|---|---|
| Pulse | Every day at 05:00 | Has a script build a snapshot from Search Console data three days back (it is typically available after 2 to 3 days, Google Search Central, 2025(opens in new tab)), runs four fixed checks, wakes every experiment due today, flags stuck ones |
| Planner | Every Monday at 06:00 | Sizes the pools, picks fix types, assigns families, creates one experiment task per arm |
| Research | 2nd of each month at 07:00 | Finds unmet demand, writes at most 8 propositions, never ships content |
| Strategy | 3rd of each month at 07:00 | Scores each fix type and proposes a memo for human approval |
The planner only uses a strategy memo whose approval task is completed. Without one, it keeps its defaults. A memo steers weeks of experiments, so a person signs it.
A pool is the set of families where one metric is measurable: AI crawler fetches, Google impressions, position or click-through rate. An arm is one fix type applied to one group of page families. A wave is the set of arms started together on one Monday, sharing one control group.
Each arm is its own task with a 10-step workflow, a graph of steps and gates, the shape we describe in What is graph engineering?.
| Step | Who runs it | Gate |
|---|---|---|
| Register | Claude Code | A real fix type and a committed record, or it stops |
| Implement | Claude Code | One variable, treated families only, product checks pass |
| Static review | Claude Code | Treated-only diff, title and meta lengths, one H1. Two rejections file a bug |
| Ship | Human | A person opens and merges the change |
| Go live (day 0) | Claude Code | Treatment verified in production HTML. Not live within 14 days: void |
| Interim (day 15) | Verdict script, read by Claude Code | Roll back on harm, void on a Google update, stop early on a strong AI-fetch win, else continue to day 31 |
| Final (day 31) | Verdict script, read by Claude Code | Win, loss or inconclusive |
| Rollback | Claude Code, then a human | Revert on a branch, a human merges, the next run confirms it is live |
| Act | Claude Code | A win files a replication idea, never a new experiment |
| Learn | Claude Code | One lesson line in the record |
Every merge waits for a person, the human gate described in Building human-in-the-loop AI systems.

Between looks, the task parks on its due date and the daily pulse wakes it.
A run, step by step (example data)
Here is one full Tidewell run, from the Monday planner to the day 31 verdict.
Monday: the planner opens a wave

The planner counts eligible families and arm capacity per pool (callout 1). Capacity is the eligible count divided by 15, rounded down, minus one: Tidewell's 68 AI-pool families allow 3 arms. The Google pool is capped at 1 arm, and each pool runs one live wave at a time.
A pool gets a wave only when its historical A/A baseline is ok. That baseline replays past data with fake splits and counts how often noise alone looks like a win (2.1% in the example). A pool that fails it gets no wave, because any verdict there would be noise.
Fix types that have won more get picked more often (callout 2), with about 20% of arms kept for fix types with no verdict (Thompson sampling, Beta(1 + wins, 1 + losses + inconclusives)).
Assignment uses the date as the random seed, shuffles within each diagnosis group, and deals families round-robin with control first (callout 3). Each group gets 17 families. The next pulse, not the planner, dispatches the four tasks.
Register: the record comes first

The record holds one fix type (callout 1), one falsifiable sentence (3), the rule that kills it (4) and a 28-day baseline (5). After the commit, only the status and the dates move. We pre-register because an agent, like a person, can find a story in any number after the fact; a committed kill rule leaves nothing to reinterpret.
Implement, review, ship, go live
On arm A, Claude Code adds a question-led summary and FAQ block to the 17 treated invoice-template families. That is the one variable. It never touches URLs, slugs, redirects, shared layout or facts, and it moves the date on every changed page. Static review passes; a human merges. Here is arm B of the same wave, parked after go-live.

Go live checks the treatment in production HTML on all 17 pages (callout 1). Google says crawling can take from a few days to a few weeks (Google Search Central, 2025(opens in new tab)), so day 0 is the day the change is verified live, not the merge day. Go live sets day 0 and both due dates, then parks the task until day 15 (callouts 2 and 3).
Day 15 and day 31: a script decides

The decision rules live in the step instructions (callout 2), so the agent cannot argue with them. It runs the verdict script and follows the matching branch.

At day 15 the script reads ratio 1.08, p 0.31: no harm, no early win, continue. At day 31 it reads ratio 1.11, p 0.19 (callout 1): inconclusive, so the task extends once to day 59 and parks again.
What the output looks like: the ledger
The ledger, not a traffic chart, is what this loop produces.

An experiment moves through registered, awaiting-merge, live, interim-done and final, and can end as extended, rolled-back or void. Each arm is its own task with its own verdict.
The verdict script runs a family-level permutation diff-in-diff over 10,000 shuffles. The ratio is the treated group's change divided by the control group's change over the same window, so 1.5 means treated moved 50% more than control. For AI fetches (each metric has its own thresholds), a win needs p ≤ 0.0477 at day 31, a ratio of 1.5 or more, and at least 60% of treated families above the control median. A day 15 AI-fetch result at p ≤ 0.0074 that also meets the ratio and share rules stops the arm early: it is final and goes straight to Act. A loss is the mirror image (ratio 0.67 or less, 60% below the median), and rollback follows. Anything else, or fewer than 15 families per arm, is inconclusive.
The two p thresholds split one 5% false-win budget across the two looks, so stopping early at day 15 costs no extra false wins. The ratio and the 60% share stop a result that is statistically real but tiny, or carried by one big page.

Inconclusive is not a loss. It says the effect, if there is one, is smaller than this site can detect. The task extends once to day 59 and then closes, because a test that runs until it wins eventually wins on noise. A win does not spawn more experiments either. It files a replication idea for a human, because a 5% false-win budget means some wins are noise, and only a second run on fresh families tells them apart.
What goes wrong: the agent that invented a crisis
Our worst morning came from an agent trying to be helpful. Our daily check skipped its check list and wrote its own report. It computed a click-through rate, compared it with an "industry standard" nobody gave it, labelled it CRITICAL, recommended changes and cited tracking bugs that had never been filed. Every number looked plausible. None came from a defined check. We archived the report and rewrote the step the same day. Here it is retold on Tidewell data.

The rule it produced has two parts. First, a script builds the snapshot, and the agent never computes a snapshot number or writes the file by hand. Second, only four fixed checks count as anomalies:
- Impressions fall below half of their 7-day mean.
- AI crawler fetches fall below 50% or rise above 300% of their 7-day mean.
- A family in a live wave is missing from the data for 7 days.
- The homepage or sitemap does not return HTTP 200 to an AI crawler.
No CTR judgements, rankings, benchmarks or recommendations. A bug is cited only by the id the tracker returned.
The second failure was quieter, and worse. A package release for one of our sites carried an unreleased product change, and it rewrote copy on one of our control pages before day 0. A control that changes is no longer a control: any gap it shows is our own edit, not the fix. We caught it before go-live. The rule: a family touched after assignment moves to "Excluded after assignment" and leaves the comparison.
Build your own: the minimal version
You do not need five workflows to start. You need one repo, two scripts, one record template and one daily scheduler.

The record template, one file per experiment:
---
experiment: exp-0001
site: example
wave: example-2026-01-05
arm: A
fix_type: internal-links
primary_metric: impressions
status: registered
branch: exp/exp-0001
day0: null
interim_due: null
final_due: null
---
Hypothesis:
One falsifiable sentence: what changes, on which treated families, against which control, by day 31.
Falsified if:
The day 31 verdict is not a win under the thresholds fixed in the verdict script.
Families:
treated (15+): ...
control (15+): ...
Baseline (28 days):
treated ... | control ...
Verdicts:
| Look | Date | n treated / control | Ratio | p | Verdict |
|---|---|---|---|---|---|
Lesson:
Copy the record template, then schedule the daily pulse in ConvOps(opens in new tab) so every experiment gets woken on its due date.
Steps
Create the ledger repo
Make one git repo with methodology, tools and a folder per site for snapshots, waves, experiments and reports. Every agent run commits its output there.
Write the snapshot script
Pull Search Console page data for three days ago, because recent days are not final yet, and write one JSON file per day. Claude Code calls the script and never computes the numbers itself.
Check what you can measure
Count the page families that have your metric. Below 15 families per group, that metric cannot carry an experiment on your site, so pick another or wait.
Assign with a seeded shuffle
Use the date as the random seed, shuffle within groups of similar pages, and deal families round-robin, control first, until each group has 15 or more.
Register before the change
Copy the record template, write one falsifiable hypothesis and the rule that kills it, and commit before any page changes.
Change treated pages and merge by hand
Have Claude Code change one variable on the treated families only, on a branch. A human merges, and day 0 is the day the change shows in production HTML.
Write the verdict script
Run a family-level permutation diff-in-diff with fixed thresholds in code: win, loss or inconclusive. The agent reads the verdict and follows the matching step.
Schedule the daily pulse
Run one job every morning that builds the snapshot, runs a few fixed checks and wakes every experiment due today, at day 15 and day 31.
Frequently asked questions
What is a pre-registered SEO experiment?
A pre-registered SEO experiment is a page change whose hypothesis, primary metric, treated and control page groups, baseline and kill rule are committed before the change ships. In this loop ConvOps runs the Register step first: Claude Code copies the record template, writes one falsifiable hypothesis and commits it to the ledger repo. Only the status and dates change afterwards, so the day 31 verdict is judged against what was promised.
Who is an SEO experiment loop for?
An SEO experiment loop is for founders and small marketing teams who run a content or product site with low traffic and make SEO changes on a hunch. ConvOps runs the loop as scheduled tasks, so each fix is kept, extended or rolled back on a controlled comparison instead of a before and after chart. The site needs at least 15 page families per group in the pool being tested, so a very small site can test only a few fix types.
Can Claude Code run SEO experiments as an agent?
Yes. Claude Code is the agent runtime, and ConvOps is the operations layer for AI agents that schedules and gates each run. Claude Code registers the experiment, edits only the treated pages on a branch, reviews the diff and checks that the change is live. It never merges the change: a human does. It never computes metrics or verdicts either. Scripts do that, and Claude Code reads the result.
How much traffic do SEO experiments need?
Enough for at least 15 page families in each group, with the metric present on every one. ConvOps runs a weekly planner that counts eligible families per pool and opens a wave only when a historical A/A baseline shows the pool is quiet enough. On a small site only large effects show up: on the AI-fetch metric a win needs a ratio of 1.5 or more, and each metric has its own thresholds in the verdict script. Click-through rate is often not measurable at all at that scale.
How long does an SEO experiment take to give a verdict?
The final verdict comes 31 days after the change is verified live. ConvOps parks the experiment task on its due date, and the daily pulse wakes it for an interim look at day 15. That look rolls back on harm, voids on a Google update, or continues to day 31, and a strong AI-fetch win at p of 0.0074 or lower stops the experiment early. An inconclusive final is extended once to day 59, then closed.
What happens when an SEO experiment loses?
A loss means the treated pages did worse than control. On the AI-fetch metric that is p of 0.0477 or lower in the harmful direction, a ratio of 0.67 or less, and at least 60% of treated families below the control median. Each metric has its own thresholds in the verdict script. ConvOps moves the task to its Rollback step, where Claude Code prepares the revert on a branch and a human merges it. The next woken run confirms the revert is live, and the loss counts against that fix type when the planner draws the next wave.
