← back to the scanner

How the score is calculated

Every number on this page is read out of the code that scores on it, so it cannot drift from what the scanner actually does. 36 crawlers across 21 vendors, five sections, and a list of the things we will not score at all.

The five sections

  • 30%answerabilityBetween two pages an assistant can read, the one whose passages survive being lifted is the one that gets quoted.
  • 25%retrievabilityA search agent that cannot fetch the page cannot cite it, and nothing else compensates.
  • 20%structureWho publishes this, and in what shape — the difference between being quoted and being attributed.
  • 15%readabilityAssistants read HTML and do not run JavaScript, so a page assembled in the browser arrives empty.
  • 10%discoveryOnly the sitemap has real adoption here, and a sitemap is table stakes rather than an advantage.

Retrievability

25% of the score

Can the agents fetch the page at all?

  • 1search12 agents whose whole job is answering a question with a citation. Blocking one costs the citation.
  • 0.6fetch8 agents that open a link a person named. Blocking one costs that visit, not your standing.
  • 0train16 agents collecting training data. Blocking them costs nothing, and we score it at zero rather than treating a publisher's deliberate choice as a fault.
  • Agents checked: 36 Across 21 vendors, each with published documentation we link to, so every verdict is checkable.

What this section refuses to score

  • An unreadable robots.txt is not an open one. A 403 or a timeout produces no score at all — the first version of this scanner gave a major newspaper 100/100 because a refusal and a wide-open file looked the same.
  • We cannot impersonate a verified crawler: the real ones are checked by reverse DNS. Where an edge refuses Googlebot alongside the assistants, that reads as a rule about unverified datacentre traffic and is reported, not scored.

Readability

15% of the score

Is there anything to read once they fetch it?

  • Served words for a full score: 250 A rendered marketing page runs to several hundred words. Chosen against real pages rather than taste.
  • Below which it is a shell: 80 A client-rendered shell serves a few dozen words of nav and footer boilerplate.

What this section refuses to score

  • The band between those two numbers is reported as partial and left out of the score. A verdict carrying 50% confidence is a coin toss, and turning a coin toss into a number is what the confidence field exists to stop.

Answerability

30% of the score

Does a passage still make sense once it is lifted off the page?

  • 0.75passagesWhether a passage lifted out of the page still makes sense on its own.
  • 0.25chunkingWhether a retriever's cuts would separate a passage from the heading it answers.
  • Words to count as an answer: 25 Shorter than this and a passage is a caption or a nav label, not something an assistant would quote as an answer.
  • Chunk sizes modelled: 1000, 2000, 4000 This is a model of somebody else's retriever, not a measurement of it, so only problems that survive every size are allowed to cost anything.

What this section refuses to score

  • A page with no passage long enough to be an answer scores null, not zero. That is a finding about the page, not a failure by it.
  • Problems that appear at one chunk size and not others are reported as fragile and cost nothing.

Structure and identity

20% of the score

Who publishes this, and is it in a shape an assistant can lift from?

  • 0.35schemaStructured data, weighted by what each type is actually consumed for.
  • 0.25liftableTables, lists and definition lists — the shapes an assistant reaches for when a question has parts.
  • 0.2entityAn Organization node with enough on it to identify you as a known entity.
  • 0.1datesWhether the declared dates are coherent. Never how recent they are.
  • 0.1titleWhether the title says which question the page answers.
  • Organization: 1 How an answer engine decides two mentions of a name are the same company. The single most useful node for being cited as an entity rather than a page.
  • Article: 0.8 Carries author and dates, which is what separates writing with a source from an undated page.
  • Product: 0.8 Names what is sold, at what price, so an answer about cost can quote you rather than a review site.
  • BreadcrumbList: 0.4 Describes where the page sits. Modest, and cheap.
  • FAQPage: 0.2 Sold hard as an AEO technique on the strength of a Google rich result Google retired. Nobody has shown an assistant treats it differently from the same text under a heading, so it is worth a little and not what anyone claims.
  • HowTo: 0.2 Same position as FAQPage: structurally tidy, no published evidence it changes an assistant's behaviour.

What this section refuses to score

  • Dates are scored for coherence, never for recency. A page that has not changed in three years may be the best answer to its question, and scoring freshness would be scoring churn.
  • A page that declares no dates at all scores exactly what a page with coherent dates scores. Absence is not a penalty.

Discovery

10% of the score

How is a crawler meant to find your other pages?

  • 1sitemapRead by every crawler for twenty years, the search-purpose AI agents included.
  • 0llmsTxtA proposal. No assistant has published that it reads /llms.txt, so scoring it would be scoring a hope. Worth adding — it costs an afternoon — but not worth a number until somebody can show it changed an outcome.
  • 0markdownTwinUnstandardised, but measurable: our own crawler logs record what share of each agent's requests asked for a markdown twin. This weight moves when that evidence does, and not before.

What this section refuses to score

  • llms.txt and markdown twins carry a weight of zero, with the reason stated: no assistant has published that it reads either. They are reported because they cost an afternoon and might matter later, and scored at nothing because today there is no evidence they do.
  • A host that refuses every sitemap probe scores null. "No sitemap" and "we were not allowed to look" are different findings and only one of them is the site's fault.

Does this score predict anything?

We do not yet know whether this score predicts citations. It is built from 0 measured sites, and we will not claim a relationship until there are at least 30. Everything above describes how readable and quotable your pages are — which is worth knowing on its own — and not how often an assistant will cite you.

0 measured sites · 30 more before we will answer at all

What we report and refuse to score

Each of these is on your report with its evidence, and none of them moves the number.

  • Evidence signals

    Every one of them says the page CLAIMS first-hand knowledge — "we surveyed", a named author, a link to a primary source. A page can carry all of them and be fabricated. Scoring them would turn "says it has data" into "has data", which is the one claim this tool exists not to make.

  • A partial render

    The verdict carries 50% confidence by design. Deciding between a shell and a page needs a rendered comparison we do not do, and turning a coin toss into a number is what the confidence field exists to stop.

  • Blocked training crawlers

    Blocking them does not affect whether you are cited. It is a publisher's deliberate choice, and scoring it would mark a decision down as a defect.

  • How recent your dates are

    A page that has not changed in three years may be the best answer to its question. We check that the dates are coherent and never that they are fresh, because scoring freshness is scoring churn.

  • A content gap found by lexical matching

    A site can answer a question in language the question does not use. A gap is where to look, and on a site we only sampled it is a statement about the pages we read rather than about the site.

What only you can measure

These are not part of the score. They need something you connect, and each one has a floor below which we report nothing rather than something shaky.

  • Search Console

    Needs: Read-only OAuth, which you can revoke at any time.

    Floor: 50 impressions before a query is reported at all.

    Google Search only. It says nothing about ChatGPT, Claude, Perplexity or Copilot, which do not report to Google.

  • Assistant referrals

    Needs: Nothing — read off your own analytics or access log.

    Floor: Shares are of attributed visits, never of all visits.

    This is a floor, not a total. Assistants strip referrers, route through redirectors, run in apps that send none, and many people read an answer then type your domain by hand. Zero here is not evidence of zero citations — it is evidence of none we could see.

  • Assistant probes

    Needs: Nothing, but it costs us model credits, so it is not run on every scan.

    Floor: 5 runs minimum, reported as a Wilson interval rather than a rate.

    Assistants are non-deterministic: the same question asked twice returns different answers and cites different pages. These rates are a sample, taken at one moment, from one model, in one place — they will not reproduce exactly, and a change between two scans is only meaningful if the intervals do not overlap.

  • Does the score predict citations

    Needs: 30 measured sites before we will answer at all.

    Floor: Spearman with a permutation test, and the honest answer printed even when it is "we do not know yet".

    An association is not a cause. Sites that score well here tend to be run by people who also do other things well, and no experiment has isolated the score from the team behind it. What this rules out is the opposite claim — a score with no relationship to citations at all.

What each part costs, and what the limit is

Every one of these is free and none of them asks for an account. That is where things stand rather than a commitment: the two that cost real resources per use — the PDF and the model probe — are the two that would change first, and the limits above are already what they cost rather than a lever to make you upgrade.

  • Scanning one page

    /tools/aeocosts the scanned site a few requests

    about a dozen requests per scan · one page at a time

    A scan is a burst of requests to somebody else's server. There is no per-scan cost to us worth counting, so there is no cap — the restraint is on how hard we hit them, not on how often you ask.

  • Comparing sites

    /tools/aeo/comparecosts the scanned site a few requests

    4 sites at once · scanned one after another, never in parallel

    Past 4 this stops being a comparison and becomes a crawl of somebody's competitor list. Sequential because four hosts hit simultaneously from one address is a burst their rate limiter is right to treat as abuse.

  • Finding content gaps

    /tools/aeo/gapscosts the scanned site a few requests

    25 pages per crawl · 10 questions · 350ms between requests

    We are a guest on your server and it should not be able to tell we were there. The page cap is also why a gap found from a small sample is worded as a place to look rather than a finding.

  • The PDF report

    /api/tools/aeo/pdfcosts us CPU and memory

    one at a time · two may wait, everyone else is asked to try again

    Each PDF launches a headless browser — a few hundred megabytes and several seconds of CPU on the box that also runs the auction. Ten at once is not ten slow responses, it is an out-of-memory kill.

  • The fix list

    /api/tools/aeo/fixescosts nothing per use

    formats a scan already taken · never scans

    It reads a stored row and writes text. There is nothing to limit.

  • The badge

    /api/tools/aeo/badgecosts nothing per use

    the score comes off after 30 days · cached for a day

    Not a quota — an expiry. A score sitting in a footer for four months is a claim nobody checked, and we serve the image.

  • Watching a page

    /tools/aeocosts the scanned site a few requests

    20 pages per address · 3 unconfirmed at a time · no more often than every 1 day

    The unconfirmed cap is the security one: without it, one person makes us deliver a hundred confirmations to somebody who never heard of us, from our sending domain.

  • Reading your access log

    /tools/aeo/logruns on your own machine or grant

    50,000 lines · parsed in your browser

    The limit is your tab's memory, not ours. An access log is a list of your visitors' IP addresses and there is no version of this report worth making us a processor of them for, so we never receive it.

  • Search Console questions

    /tools/aeo/searchruns on your own machine or grant

    your own quota · read-only, revocable from your Google account

    It runs on your grant against your data. Google's quota is the only one in play.

  • Asking a model whether it cites you

    a command, not a pagecosts us money

    5 runs minimum before a rate is reported · run by hand

    Every run costs model credits. A button for this lets a stranger spend our money by holding down a key, which is why it is a command and will stay one.

What none of these numbers can tell you

  • These are claims the page makes, not facts about it. "We surveyed 1,000 marketers" is detected whether or not anybody was surveyed — nothing here can check whether a claim is true, and a tool that pretended to would reward whoever wrote the most confident fiction.
  • This is a floor, not a total. Assistants strip referrers, route through redirectors, run in apps that send none, and many people read an answer then type your domain by hand. Zero here is not evidence of zero citations — it is evidence of none we could see.
  • Assistants are non-deterministic: the same question asked twice returns different answers and cites different pages. These rates are a sample, taken at one moment, from one model, in one place — they will not reproduce exactly, and a change between two scans is only meaningful if the intervals do not overlap.
  • An association is not a cause. Sites that score well here tend to be run by people who also do other things well, and no experiment has isolated the score from the team behind it. What this rules out is the opposite claim — a score with no relationship to citations at all.
  • Google Search only. It says nothing about ChatGPT, Claude, Perplexity or Copilot, which do not report to Google.
  • AI Overview appearances are not separated — they are folded into ordinary Search rows with no flag, so no tool can break them out from this API, including this one.
  • Data lags two to three days, and very low-volume queries are withheld by Google for privacy.