20 weeks of AI-search audits, measured.

Between 8 March and 28 July 2026, we ran a weekly automated SEO and GEO audit across eight live websites and implemented 1,683 of the fixes it produced. This is what that dataset looks like, including the parts that came back flat. The headline finding is the boring one: the technical preconditions for AI citation were solved uniformly and early across every site, so after week one they explain none of the difference between them, and the work that kept accumulating for 20 weeks was content and internal linking instead.
What exactly was measured, and how?
Every site in this dataset runs on ConceptSEO, a platform we built and operate ourselves. Once a week it runs seven analyses against each connected domain: Technical SEO, Image Audit, PageSpeed Review, Schema Audit, Content Strategy, Internal Linking, and GEO Analysis. Each analysis reads live data from Google Search Console, PageSpeed Insights, Cloudflare, and GA4, then writes individual recommendations. Every recommendation is one row with a category, a source analysis, a status, and a timestamp.
The counts below come from that ledger, queried on 28 July 2026. They are counts of work implemented, not of work suggested, so a recommendation only appears here once the change actually shipped and was verified against the live URL. The crawler measurements later in this article were taken separately with plain HTTP requests on 28 July 2026 and can be reproduced with curl in about a minute.
Eight sites is a small sample and they are not a random one. Two are our own products, one is our own marketing site, one is the platform's own site, and the rest are client sites we chose to take on. Sites are anonymised throughout and no client is named.
How much work does a weekly AI-search cadence actually generate?
Over 20 weeks, 1,683 recommendations were implemented and closed across the eight sites. 1,642 of those, or 97.6 percent, were implemented end to end by the agent rather than by a person editing code by hand.
The per-site spread is wide. Ordered by volume and stripped of names, the eight sites closed 614, 399, 280, 208, 52, 51, 50, and 29 recommendations. Most of that gap is tenure: the two largest numbers belong to the two sites connected in March, and the four smallest belong to sites connected in late June or July. Normalising for how long each site has been connected narrows it but does not close it, giving a range of roughly 9 to 31 implemented fixes per site per week.
The result we did not expect is that the rate does not decay. The two sites with 20 weeks of history are still producing at the top of that range, not the bottom. A weekly audit on a mature site does not work through a fixed backlog and go quiet. Site size is a confound here and we cannot separate it cleanly with eight sites, so read this as an observation rather than a law.
Where does the work concentrate?
Grouped by the analysis that produced them, the 1,683 implemented recommendations break down like this:
- Content Strategy: 430
- Internal Linking: 261
- GEO Analysis: 223
- Technical SEO: 220
- Schema Audit: 172
- Image Audit: 142
- PageSpeed Review: 139
- Local SEO: 29
- Weekly Review: 16
- Unlabelled (predating the source field): 51
Content Strategy and Internal Linking together account for 41 percent of everything implemented. That is the durable half of the workload. Schema, images, and PageSpeed are front-loaded: they are largely a fixed set of problems that get solved once and then only regress, which is why they sit in the middle of the table despite being the easiest work to automate.
What did not vary at all?
This is the part worth publishing even though it makes for a dull chart. Three measurements came back identical across all eight sites, which means none of them can explain any difference in outcome within this portfolio.
Fetching each homepage as GPTBot, ClaudeBot, PerplexityBot, and Google-Extended returned HTTP 200 on every one of the 32 requests. All eight sites publish a reachable llms.txt. All eight serve a real H1 and between 1,027 and 2,260 words of body text in the raw HTML to GPTBot, with a median around 1,300, so nothing here is a client-rendered shell.
Crawler access, llms.txt, and server-rendered content are the three levers most commonly written up as the way to get cited. In a portfolio where they were fixed early, they are preconditions with zero variance rather than differentiators. If you are choosing what to work on and your pages already pass those three checks, the honest answer from this dataset is that repeating them will not move anything, and the volume sits in content and internal linking.
Did the weekly cadence produce steady output?
No, and the shape is not subtle. By the month a recommendation was closed: March 400, April 118, May 490, June 150, July 525.
The analyses ran on schedule every week. What varied was the implementation, which is operator-driven and lumpy. Two of the three quiet months are the months where nobody sat down and worked the queue. Anyone reading this while pricing a subscription analysis tool should note that the audit cadence and the fix cadence are separate problems, and only the first of them is solved by software running on a timer.
What does this tell you about commissioning AI-search work?
Less than we would like, and here is what we would actually stand behind.
The technical GEO checklist is real but small. Crawler access, server-rendered HTML, schema, snippet caps, and llms.txt are worth doing and they are mostly a one-time cost. If a proposal is priced as ongoing work and its substance is that list, you are being sold a project as a retainer.
The recurring cost is content and internal linking. That is where 41 percent of the implemented work landed and it is the part that kept producing new items 20 weeks in. It is also the part that needs judgment, which is why it is the expensive half.
Generating recommendations is easy and implementing them is not. The gap between 1,831 recommendations generated and 1,683 implemented looks small here only because implementation was automated. An audit that produces 200 items for a team with no engineering time produces nothing.
What this does not show
The most important limitation is that this is not a causal study of AI citation. ConceptSEO does record which engines cited which pages for which prompts, but we are not publishing a citation-share series here, because a citation count is only as meaningful as the prompt set behind it and we have not yet run a prompt set stable enough across 20 weeks to make a before-and-after claim we would defend. Anyone quoting this article should quote it as a record of what was measured and implemented, not as evidence that a given category of change causes citations.
Beyond that: eight sites is small, they are not randomly selected, they span very different sizes and sectors, and the window is 20 weeks rather than a year. There is no control group, so nothing here separates the effect of the work from everything else that happened to these sites over the same period. The month-by-month figures reflect our own working habits as much as they reflect the sites.
The measurement window on this page closes 28 July 2026. We will republish with a longer window rather than quietly leaving these numbers to age, and if the citation-share series becomes solid enough to stand behind, that will be its own article.
If you want the method rather than the numbers, the GEO checklist we run is written up in full, and GEO vs traditional SEO covers what changes when you optimize for AI search at all.
Running this cadence on your own site is our technical SEO and GEO service, and we take on about five client builds a year.
Start a projectPhoto by Lukas Blazek on Pexels
All articles