One hundred points.Every one of them checkable.
Score state, not activity: what is true about the brand today, not what somebody is working on. Award the full points or none — partial credit is where a rubric goes to become a feeling.
Can an engine identify you with confidence, and say so without hedging?
An engine will not recommend what it cannot resolve. This is the category most brands score worst on and the one that is cheapest to fix, because almost all of it is declaration rather than production.
Layer 02 · EntityCan a machine read the page and state a fact about you without guessing?
The largest block of points, because it is the largest block of work and the one where the omissions are most consistent. On catalogs the same four fields are missing almost every time.
Layer 03 · StructureWhen a buyer asks, does your name actually appear?
The outcome variable. Everything above is a means to this, and it is the only category that cannot be improved by declaring something — it has to be measured, repeatedly, against a set you did not choose.
Layer 01 · MeasureDoes the product behave the way the marketing claims it does?
Weighted lightest, deliberately. A company can score well on visibility with no AI in its product at all, and pretending otherwise would inflate the score of anyone who bought an assistant.
Layer 04 · SystemsWill any of this still be true in twelve months?
A foundation nobody maintains decays quietly, and the decay is invisible until a competitor is named instead of you.
Layer 05 · OperateSix steps.Two people should land within a few points.
Category 03 has to be measured rather than audited, and measurement is where most self-assessments quietly cheat. Run it in this order and it is hard to flatter yourself.
Write twenty questions
The questions your buyers actually type, not the ones you wish they typed. Four each across product, brand, category, intent and comparison. Write them before you look at any result.
Pick your comparison set
Three to five competitors, named in advance. Choosing them after you see the results is how a baseline becomes a marketing document.
Run each question five times
Across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews. Private window, no account history. Five runs because the answers vary and one run is an anecdote.
Record three things per run
Which brands were named. Which domains were cited. Whether you were absent. A mention without a citation is a different result from a citation, and they should not be logged as one.
Score the rubric
Categories 01, 02, 04 and 05 are audited against the site. Category 03 is scored from the log you just produced.
Publish the number internally
Date it and keep it. A score with no prior score is a fact; two scores are a direction, and the direction is the only part anyone should act on.
The same rubric,weighted for how you are bought.
A committee and a shopper do not fail in the same place, so the weights move. Nothing is added or removed — only the point distribution changes, and both still total one hundred.
The shortlist an assistant assembles before a vendor knows it is in play.
A firm with no resolvable people is not a resolvable firm, and B2B answers lean on named expertise.
Fewer records, more depth per record. Service and FAQ schema over catalog completeness.
Unchanged, but scored on category and comparison questions rather than product ones.
The structured catalog layer an assistant reads before recommending a product.
The catalog is the asset. Completeness across every active SKU is the whole game.
A brand is easier to resolve than a firm; the product records are the harder problem.
On-site assistants matter less than being readable to the assistant the shopper already opened.
The whole rubric,as a sheet you can score by hand.
One hundred points, the six-step protocol, a twenty-question run log and the score bands, laid out to be printed and written on. No form in front of it.
Five things this standarddeliberately refuses to measure.
A rubric that scores everything scores nothing. These are the omissions, and the reasoning behind each one, so you can decide whether you agree before you use the number.
It does not measure traffic.
AI referral volume is small and rising fast, and grading yourself on it today rewards the wrong work. Presence leads volume by quarters.
It does not measure ranking.
There is no page two inside an answer. You are named or you are not, and a rubric that scores position implies a hierarchy that does not exist.
It does not predict revenue.
Nobody can honestly draw that line yet. Anyone who publishes a conversion multiplier attached to a visibility score is selling the multiplier.
It does not score content volume.
Publishing more pages is the most commonly sold remedy and one of the least reliable. Structure and entity beat volume in every baseline we have run.
It is not audited by anyone.
We wrote it, we use it, and we are publishing it so it can be argued with. Treat the weights as a considered opinion, not a finding.
Version 1.0, published August 2026. Every future revision will be numbered, dated, and will say what changed and why — including the revisions that make our own clients score lower.
