Guides

Self-Healing Website AI Agent: Heal What It Can Verify

A self-healing website AI agent is a scheduled loop that reads a site the way a reader gets it, files each defect as a bug, and has an agent in an isolated environment repair what it can rewrite and re-check, holding the rest for a person. ConvOps runs the weekly sweep, the bug pool and the per-bug heal.

DM

David Marsa

Founder & CEO

Intermediate14 min readPublished Oct 8, 2026

Last verified Oct 8, 2026

Tools and models covered:Claude CodeOpenCodeClaude Sonnet 5.5
The website as a house: a weekly sweep flags broken windows, an AI agent fixes and re-checks what it can, one waits for a person, and a shield blocks repeats.

What you will be able to do

  • Define pinned rule names and a report-only auditor that returns page, rule and detail, filed as deduplicated bugs.
  • Save the open-bug query as a pool and choose the executor that pulls from it.
  • Write an isolated environment with least privilege: one repo, tool allow-lists, own-task writes, credential references, a dispatch cap.
  • Route bugs into heal lanes and a hold lane, and verify every write with a re-lint before closing.
  • Keep a prevention registry that turns repeat findings into write-time checks.

Key takeaways

  • A self-healing website AI agent reads what a reader gets, files each defect as a bug, and repairs only what it can rewrite through the CMS API and then re-check.
  • Keep finding and fixing apart: a report-only sweep files bugs, and each bug gets its own heal run in its own isolated environment.
  • The machinery decides safety: a pool decides what gets worked, the environment decides what the agent may touch, the workflow decides when it is done.
  • A heal counts only after its re-check passes: a clean re-lint of the stored post, a moved verification stamp, or HTTP 200 with the page's own content; anything the agent cannot verify is held for a person.
  • A defect that keeps coming back belongs in a write-time refusal, not in next week's sweep.

What the self-healing website loop does

A website heals itself only where the agent can prove the fix, and we built the whole mechanism around that line. Our self-healing website AI agent sweeps the live site weekly, files every defect as a bug, and sends each bug to its own run in an isolated environment. That run rewrites what it can through the CMS API, re-checks the stored page, and closes the bug only on a clean check. ConvOps runs the schedule, the bug queue and both workflows. It is one of the loops we run our company on.

A self-healing website is a site where an agent finds reader-facing defects and repairs the ones it can verify, holding the rest for a person. It is not a self-healing test suite: those repair selectors in tests, while this loop repairs the content a reader sees.

Every screen and number below uses example data for Shutterlane, a fictional two-person camera gear review site at shutterlane.example.com.

PartWhat it is
TriggerA weekly schedule, every Tuesday at 05:00 UTC
InputsLive pages, the CMS API, the entity database behind model and directory pages
Moving partsThe sweep, the bug pool, the executor and its isolated environment, the per-bug heal workflow
OutputsBugs, heals, held bugs, prevention proposals, a run log in git
ScopeRewrites absolute-self-link, dead-link and encoding-artifact on posts. Re-checks freshness and reader-error. Holds counter-sanity, vocab-drift and any rule on a page shape it cannot write

One weekly sweep finds defects, each bug gets its own heal run, and the registry turns repeats into write-time checks.

The sweep (callout 1) files findings (2), checks them against the prevention registry (3) and dispatches open bugs (4). Each bug heals by lane (5), is verified (6), then closes or is held (7).

Why we built it: an AI agent that fixes website content bugs

A QA agent that only reports is a noise generator. An agent that finds and fixes in one run grades its own homework. We split the two.

Our sweep started report-only. It found the same broken links and stale pages every week, and nobody owned the fix. The chart shows that shape on Shutterlane data: one bug is filed (callout 1), collects a recurrence note every Tuesday (2), and is never healed (3).

A report-only sweep finds the same defect every week, and without a heal step the pile of recurrence notes only grows.

Three principles came out of it:

  1. Finding and fixing are separate jobs. The sweep never writes content, and the heal never judges the site.
  2. One defect is one unit of work. Each novel finding is its own bug, and each heal run owns exactly one bug.
  3. Done means verified. A bug closes only after a re-check of the stored page, never on the agent's word.

If your QA reports pile up the same way, see how ConvOps turns findings into bugs that executors pick up one at a time(opens in new tab).

How it runs: pool, executor, environment, workflow

The model is the least interesting part. The pool decides what gets worked, the environment what the agent can touch, the workflow when it is done. Two workflows carry the work: the sweep workflow, WF-qa, and the heal workflow, WF-bugfix. It is a graph of steps and gates, the shape we describe in What is graph engineering?, the same gated, scheduled loop as our SEO experiments with Claude Code, and the answer to agent sprawl.

A pool decides what gets worked, the executor runs it, the environment limits what it can touch, and the workflow decides when it is done.

Findings become bugs in a pool

A pool is a saved query in ConvOps: a filter over task type, tags, project, horizon and status, plus a sort. It never contains tasks; it resolves to the matching ones when asked, so bugs from any project land in every pool that matches. pools_next returns the single top task, breaks ties by task id, and records one pull row per call, so every pull is auditable.

Shutterlane's pool open-qa-bugs is type = bug AND tags contains qa-finding AND status in [backlog, todo], oldest first (callout 1).

Executors pick up work

An executor is a named runner, optionally bound to the pool it pulls from (callout 2). There are exactly two kinds:

KindHow a dispatch runs
IsolatedConvOps starts Claude Code or OpenCode, at a version from its catalog, as its own Kubernetes Job
LocalA coding agent on a machine, such as Claude Code or Kimi Code CLI, connects to ConvOps, takes the dispatch envelope and runs the task there

Our loop runs on an isolated Claude Code executor, named in the schedule. The sweep's dispatch step runs the pool-shaped query itself and sends each bug to that executor; binding the executor to a pool moves the same pull into ConvOps. A stuck run is killed at time_limit_s (per executor, 60 to 86,400 seconds, else the workspace setting: 14,400 here).

Isolated environments limit what the agent can touch

An isolated environment is the written definition of everything one run may touch. Each run is one Kubernetes Job, created suspended; a per-run Secret owned by the Job is added before it starts. A task volume keeps the task's working folder across its runs. The Job is removed 300 seconds after it finishes and the Secret goes with it, so credentials never outlive the run.

Everything the agent can touch is declared before it runs, and the credential is a reference, never a value.

  • Repos. Exactly one receives work and is always pushed (callout 1): straight onto the checked-out ref, replayed once, never forced.
  • Tool rules. The CMS server denies by default and allows 8 tools (callout 2). The browser allows read-only tools: navigate, snapshot, screenshot and three more.
  • References, not values. CMS_API_TOKEN points to a secret manager entry (callout 3), and the executor's credential is a stored reference too (callout 4). Callout 5 is the pool binding.
  • ConvOps grants. Task writes reach only the run's own task. Notes and reads reach any task, so the sweep can annotate bugs, but only a heal run closes its own bug.
  • Dispatch limits. run_dispatch_limit is 10 per run, which bounds the cost and blast radius of one sweep. run_dispatch_named_environment is false, so a run cannot widen the sandbox of what it dispatches.

Two workflows: one finds, one fixes

#Sweep stepWho runs itGate
1recall-scopeRun agentNo target host: file a red-run bug and stop
2scan-auditqa-auditor, report-onlySeven pinned rule names. A failed read is a reader-error
3file-findingsRun agentqa-dedupe on (page, rule): novel files a bug, a repeat adds a note
4preventRun agentEscape or gap: one prevention-proposal per rule
5dispatch-bugsRun agentOldest first, stop at the dispatch limit
6finalizeRun agentCommit the run log, store a memory
#Heal stepWho runs itGate
1read-bugRun agentPicks the lane. hold sets blocked and stops
2heal-postRun agentOne version-guarded write, then a clean re-lint
3heal-freshnessverifierGeneral stamp on or after the bug's creation date
4heal-readerRun agentHTTP 200 with the page's own content
5closeRun agentHealed: complete. Otherwise blocked with the reason

The auditor reads pages the way a reader gets them: some rendered in a browser, the rest as raw HTML, with counters and freshness cross-checked against the database. The freshness promise is 7 days for camera models and 14 for directory entries, the cadence their verifier loops run.

The sweep runs six steps every Tuesday at 05:00 UTC, finds and files, and never writes content.

The heal runs five steps on one bug, uses only that bug's lane, and closes it only after a re-check.

The schedule (callout 1) is FREQ=WEEKLY;BYDAY=TU;BYHOUR=5;BYMINUTE=0. Weekly matches the 7-day freshness promise, so a finding is still current when its heal runs.

A run, step by step: one bug from finding to verified fix

The interesting part of a heal is the check after the write, not the write. Here is Shutterlane's sweep qa/2026-09-15, from the schedule to the verdict.

Tuesday 05:00 UTC: the sweep finds and files

The sweep runs on isolated-sonnet in one Job with shutterlane-qa. The auditor reads 24 pages (6 in the browser, 18 as raw HTML) and returns 7 findings, one per rule.

The auditor returns a finding as page, rule and detail, and a novel page and rule pair becomes one bug titled rule: page.

This finding is novel, so it becomes bug demo0101, absolute-self-link: /blog/best-travel-cameras-2026, status todo. The run logs qa-dedupe: 5 novel, 2 recurrences; each recurrence adds one note to its open bug.

The prevent step reads the registry. absolute-self-link is covered and the post is 3 days old, inside the 7-day escape window, so the sweep files idea demo0110 as a prevention-proposal.

Dispatch lists the open QA bugs, oldest first: 7, inside the limit of 10. Each goes to isolated-sonnet without naming an environment, so it gets its project default.

One bug, one run

Each bug runs in its own isolated Job with its own volume, and notes its lane before touching anything.

demo0101 gets its own Job, run-demoex01, and its own volume. read-bug picks the lane from page shape and rule: a self-link on a /blog/ page is post. It writes bugfix-demo0101/2026-09-15 lane=post before touching anything.

One guarded write

The single write carries the version it read, so a changed post refuses it, and it never passes a status.

heal-post runs post-lint, which names the bad target and its fix. Then it makes one anchored body_md_edits replace, guarded by the expected_version it read. It never passes status, so a heal can never publish or unpublish a page.

On a version conflict the agent re-reads once and rebuilds; a second refusal ends as not healed. A dead link with no known address is unlinked and its label kept. Removing a link is safe; guessing a target is not.

Healed, or held

A bug completes only after a clean re-lint and the bad target is gone, while a hold-lane bug is blocked for a person with content untouched.

The agent re-reads and re-lints the stored post. Healed means the rule has no finding left AND the bad target is gone from the stored body, because lint can miss a target the auditor saw. The note reads healed: https://shutterlane.example.com/models/orrin-k5 -> /models/orrin-k5, the run log is committed, the bug completes, and the Job and Secret go 300 seconds later.

On the right, demo0104 (counter-sanity: /news) lands in hold. It goes blocked, "needs a product or operator fix", content untouched.

What the output looks like: report, notes, registry

Read the sweep by status, not by count. A repeat on a covered rule is a bug in your guard, not in your content.

Read the sweep by status: novel rows are new bugs, recurrences are notes, and a covered rule that recurs means the guard leaked.

ArtifactColumnsHow to read it
Sweep reportpage, rule, detail, novel or recurrence, laneNovel rows are new bugs. Recurrences are notes on open bugs
Bug note trailrun id, lane, healed: or not healed: with a reasonOne run's decision, readable without the transcript
Prevention registryrule, covered / partial / uncovered, prevented by, heal laneWhere each rule is stopped before it reaches the site

Two signals drive the prevent step:

  • Escape: a covered rule on a /news/ or /blog/ post published within 7 days. Only a recent post proves the write-time guard let it through.
  • Gap: a partial or uncovered rule on 2 or more pages in one run, or one that recurred. One page is a defect; two pages or a repeat is a pattern.

Each rule gets one proposal; an open proposal gets a note, never a second idea.

On the Shutterlane report, encoding-artifact is covered yet recurred on /directory: the guard leaking through an older publish path. We hit that pattern, and it produced a second guard, post-lint at the publish gate on top of write-time refusals in the CMS tools. Every run log lands in git, the trail we describe in AI agent observability.

What goes wrong: one rule, three spellings

Most of our failures were in the plumbing and the permissions, not the model. The worst one looked like a working loop.

Dedupe keys on (page, rule). Our auditor wrote rule names freehand and, across runs, spelled one check three ways. Dedupe saw three rules and filed three bugs for one defect. On Shutterlane data: dead-link, broken-link and dead-internal-link on /news/corvane-x2-firmware-3 became demo0081, demo0084 and demo0087.

Three spellings of one rule filed three bugs; one pinned slug leaves one bug with a recurrence note per run.

The rule it produced is one line in the auditor's card: "rule is exactly one of encoding-artifact, dead-link, absolute-self-link, counter-sanity, vocab-drift, freshness, reader-error; never a variant." Now one link is one bug with a recurrence note per run.

Smaller failureRule it produced
A run could reach only its own task, so it could not close or note other bugsGrants per environment: notes and reads on any task, tasks_dispatch and tasks_list on, task writes own-task, dispatch capped at 10
A pricing-only check did not move an entity's general freshness stampA freshness heal needs a content or status check; healed means last_verified_date moved
Dedupe matched the run's own log line and dropped every findingA corpus record identical to the incoming findings is skipped, never matched

Build your own: the minimal version

Start with the hold lane and the tool grants. Decide what the agent must never touch before you let it touch anything.

The minimal loop: a weekly schedule, a report-only auditor, a dedupe script, a bug pool and a sandboxed runner that heals and lints or holds for a person.

Keep the lint script off the network: ours judges links against the route table and entity lists, so its verdict repeats exactly. The hold lane is the human gate from Building human-in-the-loop AI systems. The lane table, ready to copy:

lanes:
  post:       # rewrite through the CMS API, then re-lint
    rules: [absolute-self-link, dead-link, encoding-artifact]
    pages: ["/blog/<slug>", "/news/<slug>"]
  freshness:  # re-verify; healed when the general stamp moves
    rules: [freshness]
    pages: ["/models/<slug>", "/directory/<slug>"]
  reader:     # re-read; healed on HTTP 200 with the page's own content
    rules: [reader-error]
  hold:       # everything else: blocked, a person fixes it
    rules: ["*"]

The environment, ready to copy:

environment: shutterlane-qa
repos:
  - repo: shutterlane/site-ops
    receives_work: true
    push_target: ref
mcp_servers:
  shutterlane-cms:
    default: deny
    allow: [post_get, post_list, post_update,
            translations_manage, model_get, tool_get,
            verification_log_manage, verdict_history]
  playwright:
    default: deny
    allow: [browser_navigate, browser_snapshot,
            browser_take_screenshot, browser_console_messages,
            browser_wait_for, browser_close]
convops_tools_any_task: [task_notes_add, task_notes_list, tasks_get]
run_dispatch_limit: 10
run_dispatch_named_environment: false
secrets:
  CMS_API_TOKEN: {kind: gcp-sm, ref: demo-project/cms-token}

Copy the lane table and the environment definition, then schedule your weekly sweep in ConvOps(opens in new tab). Want this running on your own site? Book a free Discovery Session.

Steps

  1. Pin your rule names

    List the exact rule slugs your auditor may return and forbid any variant, so deduplication on page and rule keeps working across runs.

  2. Run a report-only auditor

    Have it read pages the way a reader gets them and return one finding per defect with page, rule and detail. It never writes content.

  3. Deduplicate on page and rule

    A novel finding becomes one bug titled rule: page. A repeat adds one note to the open bug instead of a second bug.

  4. Save the open-bug query as a pool

    Filter on type bug, your QA tag and status backlog or todo, sorted oldest first, so every pull is the same deterministic query.

  5. Define the isolated environment

    One repo that receives work, the CMS MCP server with a tool allow-list, a read-only browser, task writes on the run's own task, credentials as references and a dispatch limit.

  6. Define the executor

    Pick the coding agent and its version, the model and a credential reference, then bind it to the pool or name it in the schedule.

  7. Write the heal workflow with a hold lane

    Route each bug to a lane by rule and page shape. Anything you cannot rewrite and re-check goes to hold and is blocked with its reason.

  8. Make one guarded write, then re-lint

    Write once with the version you read and never change the status. Close the bug only when the re-lint is clean and the bad target is gone.

  9. Keep a prevention registry

    Mark each rule covered, partial or uncovered, and turn escapes and repeats into write-time refusals in your CMS.

  10. Schedule the sweep weekly

    Run it on a fixed day and hour, and let its dispatch step send open bugs, oldest first, up to the dispatch limit.

Frequently asked questions

What is a self-healing website?

A self-healing website is a site where an AI agent reads pages the way a reader gets them, files each defect as a bug, repairs what it can rewrite and re-check, and holds the rest for a person. ConvOps runs the weekly sweep, the bug pool and the per-bug heal workflow. It is not a self-healing test suite, which repairs selectors inside automated tests.

Who is a self-healing website loop for?

A self-healing website loop is for small teams that run a content site through a CMS API: founders, content and marketing teams, and small engineering teams. ConvOps schedules the weekly sweep, and the sweep dispatches one heal run per bug, so broken links, encoding artifacts and stale entity pages are repaired and re-checked every week. Counter mismatches and wording drift reach a person as blocked bugs that state the reason.

How is a self-healing website different from self-healing tests?

Self-healing tests repair broken selectors inside a test suite so automated tests keep passing. A self-healing website loop in ConvOps repairs the content a reader sees on live pages: links, encoding and stale facts. It runs as scheduled workflows that write through the CMS API, re-check the stored page, and hold anything they cannot verify for a person.

How do I find and fix broken links on my website automatically?

Add the ConvOps MCP server at https://mcp.convops.app/ to your AI client and sign in once. Then create a weekly sweep workflow whose auditor returns page, rule and detail for each broken link, and file each novel finding as a bug. Dispatch each bug to a heal run that rewrites the link to its known address, or unlinks it, and re-lints the stored page before closing the bug.

Can an AI agent fix website bugs without human review?

Yes, for defects it can rewrite and verify. In ConvOps, the heal workflow rewrites three link and encoding rules on posts through the CMS API and closes a bug only after a clean re-lint. The agent never publishes or unpublishes a page. Counter mismatches and wording drift go to a hold lane, where the bug is blocked with its reason and waits for a person.

How do AI agents pick up bugs from a queue?

In ConvOps, a pool is a saved query over task type, tags, project, horizon and status that resolves to the matching tasks when asked. The pools_next call returns the single top task in the pool's sort and records the pull. An isolated or local executor runs each dispatched bug, and every heal run owns exactly one bug.

How do you sandbox an AI agent that edits your website?

ConvOps runs each isolated dispatch as its own Kubernetes Job with a per-run Secret that is removed with the Job. The environment definition names one repo that receives work, MCP servers with tool allow-lists, task writes limited to the run's own task, and a dispatch limit of 10 per run. Credentials are stored as references, never as values.

What stops an AI agent from making a wrong fix on my site?

ConvOps runs the heal as a workflow with a gate after every write. The agent makes one write per post, guarded by the version it read, and never changes the post's status. On a conflict it re-reads once, and a second refusal ends as not healed. A dead link with no known address is unlinked, not guessed, and the bug closes only on a clean re-lint.