A worksheet for the published standard. Score what is true about the brand today rather than what somebody is working on, write the date on it, and keep it — a score with no prior score is a fact, two scores are a direction, and the direction is the only part worth acting on.
The questions your buyers actually type, not the ones you wish they typed. Four each across product, brand, category, intent and comparison. Write them before you look at any result.
Three to five competitors, named in advance. Choosing them after you see the results is how a baseline becomes a marketing document.
Across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews. Private window, no account history. Five runs because the answers vary and one run is an anecdote.
Which brands were named. Which domains were cited. Whether you were absent. A mention without a citation is a different result from a citation, and they should not be logged as one.
Categories 01, 02, 04 and 05 are audited against the site. Category 03 is scored from the log you just produced.
Date it and keep it. A score with no prior score is a fact; two scores are a direction, and the direction is the only part anyone should act on.
Company: ________________________________ Scored by: ____________________ Date: ____________ Prior score: ________
| Buyer question, written before any result was seen | Type | ChatGPT | Claude | Perplexity | Gemini | AI Overviews |
|---|---|---|---|---|---|---|
| 01 | ||||||
| 02 | ||||||
| 03 | ||||||
| 04 | ||||||
| 05 | ||||||
| 06 | ||||||
| 07 | ||||||
| 08 | ||||||
| 09 | ||||||
| 10 | ||||||
| 11 | ||||||
| 12 | ||||||
| 13 | ||||||
| 14 | ||||||
| 15 | ||||||
| 16 | ||||||
| 17 | ||||||
| 18 | ||||||
| 19 | ||||||
| 20 |
Five runs per question, per engine, in a private window with no account history. Record N where the brand was named, C where an owned property was cited, and a dash where it was absent. A mention without a citation is a different result from a citation and should never be logged as one.
Can an engine identify you with confidence, and say so without hedging?
An engine will not recommend what it cannot resolve. This is the category most brands score worst on and the one that is cheapest to fix, because almost all of it is declaration rather than production.
| Criterion | Points | Your score |
|---|---|---|
| Organization schema deployed site-wide, not homepage only | 4 | |
| sameAs edges to at least four independent properties the engines already trust | 4 | |
| Named authors on published work, each with a Person node and a stable bio | 4 | |
| One canonical description of the company, identical across every property | 3 | |
| A preferred entity name stated somewhere machine-readable, with wrong variants ruled out | 2 | |
| llms.txt at the domain root, stating what the company is and how to cite it | 3 | |
| Knowledge panel present and claimed for the organization or a principal | 3 | |
| No contradictory facts across owned properties: one headcount, one founding date, one title per person | 2 | |
| Subtotal — entity resolution | 25 |
Can a machine read the page and state a fact about you without guessing?
The largest block of points, because it is the largest block of work and the one where the omissions are most consistent. On catalogs the same four fields are missing almost every time.
| Criterion | Points | Your score |
|---|---|---|
| Server-rendered HTML: the content exists without executing JavaScript | 6 | |
| Product or Service schema across the full active catalog or service set | 5 | |
| Variant-level availability and priceCurrency present where a price is shown | 3 | |
| shippingDetails present | 2 | |
| returnPolicy present | 2 | |
| FAQPage schema on priority pages, written in the brand's own voice | 3 | |
| Hub pages that answer the comparison questions buyers actually ask | 4 | |
| Answer-shaped copy: the page states the claim, then supports it, rather than burying it | 3 | |
| Passes Rich Results and Schema.org validation with no errors | 2 | |
| Subtotal — structure | 30 |
When a buyer asks, does your name actually appear?
The outcome variable. Everything above is a means to this, and it is the only category that cannot be improved by declaring something — it has to be measured, repeatedly, against a set you did not choose.
| Criterion | Points | Your score |
|---|---|---|
| A baseline exists: twenty real buyer questions, five runs each, five engines, recorded | 5 | |
| Brand named in at least three of ten runs on the primary category question | 6 | |
| Owned properties cited, not just mentioned, in at least one engine | 5 | |
| Present in at least three of the five engines | 4 | |
| Measured against three to five named competitors, not against nobody | 3 | |
| Coverage across all four query types: product, brand, category and intent | 2 | |
| Subtotal — answer presence | 25 |
Does the product behave the way the marketing claims it does?
Weighted lightest, deliberately. A company can score well on visibility with no AI in its product at all, and pretending otherwise would inflate the score of anyone who bought an assistant.
| Criterion | Points | Your score |
|---|---|---|
| Any customer-facing AI runs retrieval over a first-party corpus, not a prompt over a general model | 3 | |
| A human review gate before anything generated publishes | 3 | |
| Cost, latency and failure rate measured per call | 2 | |
| An evaluation harness, so a model change is caught before a customer finds it | 2 | |
| Subtotal — systems | 10 |
Will any of this still be true in twelve months?
A foundation nobody maintains decays quietly, and the decay is invisible until a competitor is named instead of you.
| Criterion | Points | Your score |
|---|---|---|
| Citation share re-read on a fixed cadence, at least quarterly | 4 | |
| A named owner internally, not an agency retainer line | 2 | |
| Handoff documentation covering schema, content systems and identity sources | 2 | |
| Alerting or a scheduled check when a model update moves a tracked answer | 2 | |
| Subtotal — operation | 10 |
Categories 01, 02, 04 and 05 are audited against the site. Category 03 is scored from the run log above.
Resolvable, readable and present. Defending a position rather than building one.
Readable, inconsistently named. Usually an answer-presence problem sitting on a sound structure.
Structurally legible in places. The common score for a well-built site that was never built for machines.
Not resolvable. The engines are answering your category without you, and no amount of content fixes it until the entity does.
The shortlist an assistant assembles before a vendor knows it is in play.
The structured catalog layer an assistant reads before recommending a product.
Use the variant that matches how you sell. The protocol does not change; only the weights do, and every reweighting is stated rather than applied quietly.
AI referral volume is small and rising fast, and grading yourself on it today rewards the wrong work. Presence leads volume by quarters.
There is no page two inside an answer. You are named or you are not, and a rubric that scores position implies a hierarchy that does not exist.
Nobody can honestly draw that line yet. Anyone who publishes a conversion multiplier attached to a visibility score is selling the multiplier.
Publishing more pages is the most commonly sold remedy and one of the least reliable. Structure and entity beat volume in every baseline we have run.
We wrote it, we use it, and we are publishing it so it can be argued with. Treat the weights as a considered opinion, not a finding.
We wrote this, we use it, and we publish it so it can be argued with. Treat the weights as a considered opinion rather than a finding. Full standard, with the reasoning behind each weight: pyxl.com/airo