Skip to Content
The standard

AIRO v1.0. Read it, run it, use it without us.

AIRO is how we measure whether an answer engine can find, read and recommend a company. This is the whole rubric: 100 points, five categories, and a protocol specific enough that two people running it on the same brand should land within a few points of each other. It is free, it requires no relationship with us, and if your score comes back fine then you do not need us.

100
Points, across five categories
5
Engines every score is read against
v1.0
Published August 2026, versioned from here
$0
To run it yourself
Why publish it

A named method nobody can readis just a trademark.

Every firm in this category has a proprietary framework with a name and no contents. You cannot evaluate one, so you end up evaluating the pitch instead. We would rather be checked.

Publishing the rubric costs us the ability to be mysterious about it. That is the point: it is only credible to say we are not afraid of being copied if we behave that way. If you run this and fix what it finds without ever calling us, the standard did its job.

85–100Resolvable, readable and present. Defending a position rather than building one.
65–84Readable, inconsistently named. Usually an answer-presence problem sitting on a sound structure.
40–64Structurally legible in places. The common score for a well-built site that was never built for machines.
0–39Not resolvable. The engines are answering your category without you, and no amount of content fixes it until the entity does.

Most brands we baseline land between 40 and 64. Almost none of them are badly built. They were built for one reader.

The rubric

One hundred points.Every one of them checkable.

Score state, not activity: what is true about the brand today, not what somebody is working on. Award the full points or none — partial credit is where a rubric goes to become a feeling.

01Entity resolution
25points

Can an engine identify you with confidence, and say so without hedging?

An engine will not recommend what it cannot resolve. This is the category most brands score worst on and the one that is cheapest to fix, because almost all of it is declaration rather than production.

Layer 02 · Entity
Organization schema deployed site-wide, not homepage only4
sameAs edges to at least four independent properties the engines already trust4
Named authors on published work, each with a Person node and a stable bio4
One canonical description of the company, identical across every property3
A preferred entity name stated somewhere machine-readable, with wrong variants ruled out2
llms.txt at the domain root, stating what the company is and how to cite it3
Knowledge panel present and claimed for the organization or a principal3
No contradictory facts across owned properties: one headcount, one founding date, one title per person2
02Structure
30points

Can a machine read the page and state a fact about you without guessing?

The largest block of points, because it is the largest block of work and the one where the omissions are most consistent. On catalogs the same four fields are missing almost every time.

Layer 03 · Structure
Server-rendered HTML: the content exists without executing JavaScript6
Product or Service schema across the full active catalog or service set5
Variant-level availability and priceCurrency present where a price is shown3
shippingDetails present2
returnPolicy present2
FAQPage schema on priority pages, written in the brand's own voice3
Hub pages that answer the comparison questions buyers actually ask4
Answer-shaped copy: the page states the claim, then supports it, rather than burying it3
Passes Rich Results and Schema.org validation with no errors2
03Answer presence
25points

When a buyer asks, does your name actually appear?

The outcome variable. Everything above is a means to this, and it is the only category that cannot be improved by declaring something — it has to be measured, repeatedly, against a set you did not choose.

Layer 01 · Measure
A baseline exists: twenty real buyer questions, five runs each, five engines, recorded5
Brand named in at least three of ten runs on the primary category question6
Owned properties cited, not just mentioned, in at least one engine5
Present in at least three of the five engines4
Measured against three to five named competitors, not against nobody3
Coverage across all four query types: product, brand, category and intent2
04Systems
10points

Does the product behave the way the marketing claims it does?

Weighted lightest, deliberately. A company can score well on visibility with no AI in its product at all, and pretending otherwise would inflate the score of anyone who bought an assistant.

Layer 04 · Systems
Any customer-facing AI runs retrieval over a first-party corpus, not a prompt over a general model3
A human review gate before anything generated publishes3
Cost, latency and failure rate measured per call2
An evaluation harness, so a model change is caught before a customer finds it2
05Operation
10points

Will any of this still be true in twelve months?

A foundation nobody maintains decays quietly, and the decay is invisible until a competitor is named instead of you.

Layer 05 · Operate
Citation share re-read on a fixed cadence, at least quarterly4
A named owner internally, not an agency retainer line2
Handoff documentation covering schema, content systems and identity sources2
Alerting or a scheduled check when a model update moves a tracked answer2
How to run it

Six steps.Two people should land within a few points.

Category 03 has to be measured rather than audited, and measurement is where most self-assessments quietly cheat. Run it in this order and it is hard to flatter yourself.

01

Write twenty questions

The questions your buyers actually type, not the ones you wish they typed. Four each across product, brand, category, intent and comparison. Write them before you look at any result.

02

Pick your comparison set

Three to five competitors, named in advance. Choosing them after you see the results is how a baseline becomes a marketing document.

03

Run each question five times

Across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews. Private window, no account history. Five runs because the answers vary and one run is an anecdote.

04

Record three things per run

Which brands were named. Which domains were cited. Whether you were absent. A mention without a citation is a different result from a citation, and they should not be logged as one.

05

Score the rubric

Categories 01, 02, 04 and 05 are audited against the site. Category 03 is scored from the log you just produced.

06

Publish the number internally

Date it and keep it. A score with no prior score is a fact; two scores are a direction, and the direction is the only part anyone should act on.

Two variants

The same rubric,weighted for how you are bought.

A committee and a shopper do not fail in the same place, so the weights move. Nothing is added or removed — only the point distribution changes, and both still total one hundred.

AIRO for B2B

The shortlist an assistant assembles before a vendor knows it is in play.

Entity resolution25 → 30

A firm with no resolvable people is not a resolvable firm, and B2B answers lean on named expertise.

Structure30 → 25

Fewer records, more depth per record. Service and FAQ schema over catalog completeness.

Answer presence25 → 25

Unchanged, but scored on category and comparison questions rather than product ones.

AIRO Commerce

The structured catalog layer an assistant reads before recommending a product.

Structure30 → 38

The catalog is the asset. Completeness across every active SKU is the whole game.

Entity resolution25 → 20

A brand is easier to resolve than a firm; the product records are the harder problem.

Systems10 → 7

On-site assistants matter less than being readable to the assistant the shopper already opened.

06Take it with you

The whole rubric,as a sheet you can score by hand.

One hundred points, the six-step protocol, a twenty-question run log and the score bands, laid out to be printed and written on. No form in front of it.

Printable, letter or A4AIRO v1.0 — the scorecardEvery criterion with its point value and a box to score it in, a run log for the twenty questions across all five engines, both variant reweightings, and the five things the standard deliberately does not measure.100Points to score20Log rows5Engine columnsOpen it, then print or save as PDF ⟶

Tell us where to send v1.1.Nothing else.

The weights will move as the engines do, and every change will be stated on the record rather than applied quietly. If you have scored yourself against v1.0, the revision is the only thing worth knowing about.

Your address goes into our own CRM and nowhere else. Privacy policy.

The limits, stated

Five things this standarddeliberately refuses to measure.

A rubric that scores everything scores nothing. These are the omissions, and the reasoning behind each one, so you can decide whether you agree before you use the number.

It does not measure traffic.

AI referral volume is small and rising fast, and grading yourself on it today rewards the wrong work. Presence leads volume by quarters.

It does not measure ranking.

There is no page two inside an answer. You are named or you are not, and a rubric that scores position implies a hierarchy that does not exist.

It does not predict revenue.

Nobody can honestly draw that line yet. Anyone who publishes a conversion multiplier attached to a visibility score is selling the multiplier.

It does not score content volume.

Publishing more pages is the most commonly sold remedy and one of the least reliable. Structure and entity beat volume in every baseline we have run.

It is not audited by anyone.

We wrote it, we use it, and we are publishing it so it can be argued with. Treat the weights as a considered opinion, not a finding.

Version 1.0, published August 2026. Every future revision will be numbered, dated, and will say what changed and why — including the revisions that make our own clients score lower.

Run it yourself.Or have us run it against your category.

If your score comes back fine, you do not need us and we will say so. If it does not, we will put twenty real buyer questions through all five engines and show you every citation, every mention and every miss against three to five competitors.

We would like to measure which pages people read, using Google Analytics. That is the only thing these cookies do here — no advertising, no profile sold to anyone, nothing shared outside Pyxl. Say no and the site works exactly the same. Privacy policy.