The AIRO v1.0 rubric as a worksheet. Print it, or save it as a PDF.Read the standard in full ⟶
AIRO v1.0 · the scorecard · pyxl.com/airo-scorecardFree to use, with or without us

AIRO v1.0 the scorecard

Published August 2026
Versioned from here
pyxl.com/airo

A worksheet for the published standard. Score what is true about the brand today rather than what somebody is working on, write the date on it, and keep it — a score with no prior score is a fact, two scores are a direction, and the direction is the only part worth acting on.

100Points, five categories
5Engines every score is read against
20Buyer questions in the run log
$0To run it yourself
Before you score · the protocol
01Write twenty questions

The questions your buyers actually type, not the ones you wish they typed. Four each across product, brand, category, intent and comparison. Write them before you look at any result.

02Pick your comparison set

Three to five competitors, named in advance. Choosing them after you see the results is how a baseline becomes a marketing document.

03Run each question five times

Across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews. Private window, no account history. Five runs because the answers vary and one run is an anecdote.

04Record three things per run

Which brands were named. Which domains were cited. Whether you were absent. A mention without a citation is a different result from a citation, and they should not be logged as one.

05Score the rubric

Categories 01, 02, 04 and 05 are audited against the site. Category 03 is scored from the log you just produced.

06Publish the number internally

Date it and keep it. A score with no prior score is a fact; two scores are a direction, and the direction is the only part anyone should act on.

Company: ________________________________   Scored by: ____________________   Date: ____________   Prior score: ________

The run log · category 03 is scored from this
Buyer question, written before any result was seenTypeChatGPTClaudePerplexityGeminiAI Overviews
01
02
03
04
05
06
07
08
09
10
11
12
13
14
15
16
17
18
19
20

Five runs per question, per engine, in a private window with no account history. Record N where the brand was named, C where an owned property was cited, and a dash where it was absent. A mention without a citation is a different result from a citation and should never be logged as one.

01

Entity resolution

25 pts

Can an engine identify you with confidence, and say so without hedging?

An engine will not recommend what it cannot resolve. This is the category most brands score worst on and the one that is cheapest to fix, because almost all of it is declaration rather than production.

CriterionPointsYour score
Organization schema deployed site-wide, not homepage only4
sameAs edges to at least four independent properties the engines already trust4
Named authors on published work, each with a Person node and a stable bio4
One canonical description of the company, identical across every property3
A preferred entity name stated somewhere machine-readable, with wrong variants ruled out2
llms.txt at the domain root, stating what the company is and how to cite it3
Knowledge panel present and claimed for the organization or a principal3
No contradictory facts across owned properties: one headcount, one founding date, one title per person2
Subtotal — entity resolution25
02

Structure

30 pts

Can a machine read the page and state a fact about you without guessing?

The largest block of points, because it is the largest block of work and the one where the omissions are most consistent. On catalogs the same four fields are missing almost every time.

CriterionPointsYour score
Server-rendered HTML: the content exists without executing JavaScript6
Product or Service schema across the full active catalog or service set5
Variant-level availability and priceCurrency present where a price is shown3
shippingDetails present2
returnPolicy present2
FAQPage schema on priority pages, written in the brand's own voice3
Hub pages that answer the comparison questions buyers actually ask4
Answer-shaped copy: the page states the claim, then supports it, rather than burying it3
Passes Rich Results and Schema.org validation with no errors2
Subtotal — structure30
03

Answer presence

25 pts

When a buyer asks, does your name actually appear?

The outcome variable. Everything above is a means to this, and it is the only category that cannot be improved by declaring something — it has to be measured, repeatedly, against a set you did not choose.

CriterionPointsYour score
A baseline exists: twenty real buyer questions, five runs each, five engines, recorded5
Brand named in at least three of ten runs on the primary category question6
Owned properties cited, not just mentioned, in at least one engine5
Present in at least three of the five engines4
Measured against three to five named competitors, not against nobody3
Coverage across all four query types: product, brand, category and intent2
Subtotal — answer presence25
04

Systems

10 pts

Does the product behave the way the marketing claims it does?

Weighted lightest, deliberately. A company can score well on visibility with no AI in its product at all, and pretending otherwise would inflate the score of anyone who bought an assistant.

CriterionPointsYour score
Any customer-facing AI runs retrieval over a first-party corpus, not a prompt over a general model3
A human review gate before anything generated publishes3
Cost, latency and failure rate measured per call2
An evaluation harness, so a model change is caught before a customer finds it2
Subtotal — systems10
05

Operation

10 pts

Will any of this still be true in twelve months?

A foundation nobody maintains decays quietly, and the decay is invisible until a competitor is named instead of you.

CriterionPointsYour score
Citation share re-read on a fixed cadence, at least quarterly4
A named owner internally, not an agency retainer line2
Handoff documentation covering schema, content systems and identity sources2
Alerting or a scheduled check when a model update moves a tracked answer2
Subtotal — operation10
Total score

Categories 01, 02, 04 and 05 are audited against the site. Category 03 is scored from the run log above.

/ 100
What the number means
85–100

Resolvable, readable and present. Defending a position rather than building one.

65–84

Readable, inconsistently named. Usually an answer-presence problem sitting on a sound structure.

40–64

Structurally legible in places. The common score for a well-built site that was never built for machines.

0–39

Not resolvable. The engines are answering your category without you, and no amount of content fixes it until the entity does.

Two variants · same protocol, different weights

AIRO for B2B

The shortlist an assistant assembles before a vendor knows it is in play.

Entity resolution25 → 30
Structure30 → 25
Answer presence25 → 25

AIRO Commerce

The structured catalog layer an assistant reads before recommending a product.

Structure30 → 38
Entity resolution25 → 20
Systems10 → 7

Use the variant that matches how you sell. The protocol does not change; only the weights do, and every reweighting is stated rather than applied quietly.

What this does not measure
It does not measure traffic.

AI referral volume is small and rising fast, and grading yourself on it today rewards the wrong work. Presence leads volume by quarters.

It does not measure ranking.

There is no page two inside an answer. You are named or you are not, and a rubric that scores position implies a hierarchy that does not exist.

It does not predict revenue.

Nobody can honestly draw that line yet. Anyone who publishes a conversion multiplier attached to a visibility score is selling the multiplier.

It does not score content volume.

Publishing more pages is the most commonly sold remedy and one of the least reliable. Structure and entity beat volume in every baseline we have run.

It is not audited by anyone.

We wrote it, we use it, and we are publishing it so it can be argued with. Treat the weights as a considered opinion, not a finding.

We wrote this, we use it, and we publish it so it can be argued with. Treat the weights as a considered opinion rather than a finding. Full standard, with the reasoning behind each weight: pyxl.com/airo