Home/Blog/AEO audit

How to Audit Your Site for AEO: A 25-Point Readiness Scorecard

A replicable scoring method for whether your site is positioned to be cited by AI answer engines. Run it yourself in about ninety minutes.

Updated August 2026
How to Audit Your Site for AEO: A 25-Point Readiness Scorecard
The short answer

Score your site across five layers, five checks each: retrievability, extractability, corroboration, entity clarity, and measurement. Award one point per check passed. Below 10 means fix foundations before anything else. 10 to 17 means you are structurally sound and need off-site work. Above 17 means measure and refine.

Most SEO audits I see, including good ones, do not test the things that decide whether an AI system will quote you. They test crawlability, speed, metadata, and link profile, all of which matter and none of which are sufficient.

So I built a scorecard for the gap. It is deliberately simple, deliberately manual, and deliberately replicable, which matters because a scoring method nobody else can run is an opinion rather than a benchmark.

Everything below can be done without paid tools. Score honestly. A generous score produces a comfortable document and no improvement.

How to use the scorecard

Award one point per check that clearly passes. Partial credit defeats the purpose. If you are arguing with yourself about whether something counts, it does not count.

Test against your five most commercially important pages rather than your whole site. If those five fail, the rest almost certainly do, and if they pass, you have a template to apply elsewhere.

Write the date on it. The value is in re-running it quarterly and watching the number move, not in the first score.

Section 1: Retrievability (5 points)

Can AI systems reach and read the content at all. Everything else is theoretical until this passes.

1. Content renders without JavaScript. Disable JavaScript in your browser and load the page. If the main text and internal links vanish, you fail. This is the single most common hard failure I find.

2. Internal links exist in the raw HTML. View source and search for your key links. Links injected client-side may not be followed.

3. robots.txt does not block major AI crawlers unless you have deliberately chosen to. Check for GPTBot, ClaudeBot, PerplexityBot, and Google-Extended. Choosing to block is legitimate. Blocking by accident is not.

4. The page returns HTTP 200 and has one canonical URL. No redirect chains, no duplicate versions competing.

5. The page is indexed. Search for a distinctive exact phrase from it in quotation marks. If nothing returns, nothing downstream matters.

How to Audit Your Site for AEO: A 25-Point Readiness Scorecard

Section 2: Extractability (5 points)

Whether a model can lift a clean, quotable answer. This is where the published evidence is strongest and where most sites lose points.

6. An answer capsule appears in the first 30% of the page. A 40 to 60 word direct answer to the page's core question, self-contained enough to quote without surrounding context. Roughly 44.2% of LLM citations come from that first 30%, so placement is not cosmetic.

7. Headings appear every 120 to 180 words. Count the words between your H2s and H3s on one page. That density correlates with about 70% more citations than sparse structure.

8. The page contains at least one table or structured list where it genuinely aids comprehension. Explicitly structured content is cited roughly two to three times more often than pure prose.

9. An explicit FAQ block exists with questions as headings and self-contained answers beneath them.

10. The page contains at least one claim that is specifically yours. A number, a framework, an observation from your own work. The strongest citation configuration pairs an answer capsule with proprietary insight, at roughly a 34.3% citation rate. Restating consensus scores you nothing here.

Section 3: Corroboration (5 points)

Whether anything outside your own website confirms that you exist and know this subject.

11. Your business is mentioned by name on at least five domains you do not control. Directories count. Reprints of your own press release do not.

12. At least one of those mentions is editorial rather than a listing. Someone chose to reference you.

13. You appear in the recognized directories for your industry. For agencies that means the likes of Clutch or GoodFirms. For professional services it means the relevant bar, institute, or association listings.

14. Your key people have public profiles with consistent titles and a link back to the site.

15. You have been referenced in a discussion venue where your topic is actually debated. Forums, communities, comment threads, industry groups. This is the hardest point on the scorecard and the most valuable, particularly if citation really is partly post-hoc.

Section 4: Entity clarity (5 points)

Whether machines can resolve who you are without ambiguity.

16. Organization schema exists with name, URL, logo, and contact details.

17. sameAs links connect your site to your external profiles. LinkedIn, industry directories, social accounts.

18. Your business name, address, and phone number are identical everywhere. Character for character. Abbreviation differences count as failures.

19. Your description of what you do is consistent across site, profiles, and listings. Three different self-descriptions makes resolution harder for no benefit.

20. Author identity is machine-readable on editorial content. Person schema with jobTitle and sameAs, applied consistently. A site whose articles have no resolvable author is harder to attribute confidently.

Section 5: Measurement (5 points)

Whether you can tell if any of this is working.

21. You have a fixed prompt set of ten to twenty questions a buyer might ask an assistant, written down and unchanged between runs.

22. You run that set at least monthly against the main assistants and record whether you were mentioned.

23. You can identify referral traffic from AI sources in analytics.

24. You track branded search volume over time. If off-site presence is working, people begin searching your name, and that is often the earliest reliable signal.

25. You have a baseline from at least ninety days ago to compare against. Without it you are measuring noise.

Interpreting your score

ScoreWhat it meansWhat to do next
0–9Foundations are brokenFix Sections 1 and 2 only. Ignore everything else until you are above 10
10–17Structurally sound, under-corroboratedThe bottleneck is almost certainly Section 3. On-page tuning will not move it
18–22Genuinely well positionedRefine, measure, and go deeper on one topic rather than broader
23–25Either excellent or scored generouslyRe-score with someone else doing it

The distribution I see in practice clusters heavily in the 8 to 14 band, and the pattern is consistent: reasonable Section 1, weak Section 2, very weak Section 3, accidental Section 4, absent Section 5.

Why most sites fail Section 3 hardest

Worth naming, because it is where the work actually is.

Sections 1, 2, and 4 are all things you can fix on your own website in a few weeks. They are within your control, they are cheap, and a competent practitioner can complete them quickly.

Section 3 is other people deciding you are worth mentioning. It cannot be bought without buying something worthless, it cannot be rushed, and it is the layer the 2026 research suggests matters most. That asymmetry, cheap controllable work versus expensive uncontrollable work, is why so much AEO advice concentrates on the first kind. It is easier to sell.

If your score is 14 with a Section 3 of 1, more on-page optimization is not your answer, and anyone selling it to you as one is selling the easy half.

What to fix first, in order

Work upward through the layers rather than picking the interesting problems.

Any Section 1 failure comes first, always, because everything else is wasted underneath it. Then Section 2 on your five most important pages, which is usually a week of focused rewriting and delivers the fastest visible change. Then Section 4, which is roughly a day and is embarrassing to leave undone given how cheap it is. Then Section 5, so you can see what happens next. Then Section 3, continuously, forever, because it never finishes.

Three worked scores

Scoring bands mean more with examples attached, so here are three composite profiles drawn from the patterns I see most often. None is a real client.

The well-built invisible site: 13 out of 25. A design agency rebuilt its site eighteen months ago on a modern framework. Section 1 scores 3, losing points because the case study content is injected client-side and the internal links to it do not exist in the raw HTML. Section 2 scores 2, because the pages are beautiful, prose-heavy, and contain nothing quotable in the first third. Section 3 scores 3, mostly from directory listings. Section 4 scores 4, since the developer implemented schema properly. Section 5 scores 1. The instinct here is always to improve the content further. The actual fix is the rendering problem in Section 1, which is invalidating everything above it, and then the extraction work in Section 2. Both are engineering tasks, not writing tasks.

The prolific publisher: 11 out of 25. A consultancy has published weekly for four years, so it has around two hundred posts. Section 1 scores 5, since it is a plain content management system that works fine. Section 2 scores 1, because two hundred posts all follow the same house style of long introductions and no structure. Section 3 scores 2. Section 4 scores 2. Section 5 scores 1. This is the most frustrating profile, because enormous effort has already been spent and it is not converting into citability. The answer is emphatically not more posts. It is consolidating the two hundred into perhaps forty genuinely good ones, restructured, with the rest redirected. That recommendation is usually unwelcome.

The known name with a bad site: 16 out of 25. A regional firm whose founder speaks at industry events and is quoted in trade press. Section 1 scores 4, Section 2 scores 2, Section 3 scores 5, which is rare, Section 4 scores 4, Section 5 scores 1. This firm gets mentioned in AI answers already, inconsistently, despite the site being mediocre, which is exactly what the post-hoc citation hypothesis would predict. Their upside is unusually cheap: a week of Section 2 work converts existing recognition into reliable citation. They are the only one of the three where on-page work alone will move the number meaningfully.

The pattern worth extracting is that the same score can mean completely different things depending on where the points are. A 13 concentrated in Sections 3 and 4 is a much better position than a 13 concentrated in Sections 1 and 2, because the expensive, slow work is already done. Always read the section breakdown rather than the total.

Frequently asked questions

How often should I re-run this audit?

Quarterly. Sections 1, 2, and 4 change only when you change them. Sections 3 and 5 accumulate slowly, and a quarter is roughly the interval at which movement becomes visible above the noise.

Can I automate the scoring?

Sections 1 and 4 largely yes, with standard crawling and schema validation tools. Sections 2 and 3 need human judgment, since the questions are about quality and genuine independence rather than presence. Section 5 is about your own process. I would keep it manual, because the argument you have with yourself while scoring is where most of the value is.

Is a low score bad news?

It is information. Most sites score low, and the two sections that produce the fastest improvement, 1 and 2, are also the cheapest to fix. A score of 8 with a clear path to 15 is a better position than a score of 15 with no idea what to do next.

Does this replace a traditional SEO audit?

No. It tests a different axis. A site can pass this scorecard and still be invisible in classic search because of competitive weakness, and the reverse is also true. Run both.

Why 25 points rather than a weighted score?

Because weighting requires knowing the relative importance of each factor, and the evidence is not strong enough to justify specific weights yet. An unweighted count is honest about that uncertainty. If better data emerges, weighting is the obvious next version.

The value of a scorecard is not the number. It is that it forces you to look at the layer you have been avoiding, which for most sites is the one that requires other people to care.

If you would rather have someone else score it, send me your five most important URLs and I will run it and send back the marked scorecard. ADPH is a US-facing consultancy delivering its work from the Philippines, and the audit is the same either way.