Share of Answer is how often an AI assistant names a business across a defined, frozen set of buying questions, on named platforms, over a stated number of runs. It is a measure of presence, never of rank, and never more precise than its sample allows.
- There is no such thing as an AI ranking. Ask the same question a hundred times and you will get a different list, in a different order, nearly every time. The research is below and it is not close.
- So we measure presence, repeatedly. How often you are named, out of a set of questions we froze in advance.
- Two instruments, one metric. A monthly 12-question screen reported as a fraction, and a quarterly ≥400-observation measurement reported as a percentage with a confidence interval. The unit changes because the sample size changes.
- Movement has a threshold. Below it, we do not narrate the number as progress, even when it moved our way.
- Everything here is public on purpose. A measurement standard nobody can inspect is not a standard.
Why “AI rank tracking” is selling you a coordinate that does not exist
In January 2026 Rand Fishkin published research with Patrick O’Donnell of Gumshoe.ai: 600 volunteers ran 12 prompts across ChatGPT, Claude and Google’s AI Overview, producing 2,961 responses. The finding, quoted exactly: “there’s a <1 in 100 chance that ChatGPT or Google’s AI, if asked 100X, will give you the same list of brands in any two responses”, and on ordering, “it’s more like 1 in 1,000 runs before you’d see two lists in the same order.”
Ask an AI for recommendations a hundred times and nearly every response differs in three ways at once: the brands listed, the order they appear in, and how many there are.
The same instability shows up on Google’s side. Ahrefs tracked 43,000 keywords with at least sixteen recorded AI Overviews each over a month and found AI Overviews have “a 70% chance of changing from one observation to the next”, “a persistence of 2.15 days on average”, and that between consecutive responses only 54.5% of cited URLs overlap. Nearly half the sources are entirely new each time.
Read those together and one conclusion follows. A product showing you “your rank in ChatGPT” is showing you a coincidence with a number printed on it. A weekly AI visibility check is measuring weather, not climate. The only thing stable enough to track is how often you appear at all, sampled across enough runs to mean something.
One metric, two instruments, and why the unit changes
Precision comes from observation count. That single fact governs everything below, and it is why we report Share of Answer two different ways rather than pretending one number fits both jobs.
A twelve-question screen cannot carry a percentage. With 108 observations behind it, “58%” implies a resolution the sample does not have. So the screen stays a fraction, X out of 12, which keeps the sample size visible inside the number itself. 7/12 tells you the result and the evidence; a percentage tells you the result and hides the evidence.
A quarterly run of 400-plus observations can. At roughly 400 observations the 95% confidence interval is about ±5 percentage points, which is enough to state a percentage honestly, provided the interval is printed next to it, every time. At around 100 observations it is nearer ±10pp, which is why a per-platform slice is never given the blended figure’s precision.
So: the free monthly screen is a fraction. The quarterly measurement is a percentage with its interval attached. Same metric, same house term, two instruments, and any number we publish says which one produced it.
| Tier 1: the screen | Tier 2: the measurement | |
|---|---|---|
| Cadence | Monthly | Quarterly |
| Question set | 12 buying questions, frozen and versioned | 10–15 money queries per topic-space, plus fan-out variants |
| Platforms | Google AI Overviews · ChatGPT · Perplexity | The same three, plus AI Mode as spot coverage |
| Runs | 3 per question per platform | Enough to reach ≥400 observations per topic-space |
| Observations | ~108 | ≥400 |
| Reported as | X / 12 per platform, as a fraction | Visibility-% blended, ±5pp, with the interval printed |
| Movement rule | Read as trend only; single months are not narrated | |Δ| ≥ 5pp = movement · 2.5–5pp = watch · <2.5pp = noise |
| What it is for | Triage and direction: is anything structurally broken? | Evidence: did the programme move the needle? |
The 12-question protocol, in full
This is the method behind the free Visibility Check, and behind the X/12 figures quoted everywhere else on this site. Nothing here is proprietary. Run it yourself if you would rather not take our word for anything.
- 01Build twelve real buying questionsQuestions a customer asks before They know your name: “who should replace a privacy fence in Marietta,” not “is {your company} any good.” Brand queries inflate the score and prove only that you exist. Vary the phrasing deliberately: people asking the same thing word it in a dozen ways, so one canonical keyword measures almost nothing.
- 02Freeze the set and version itWrite the twelve down and do not change them between runs. A changed question set resets the baseline, and comparing across versions is not a comparison. If you must revise, note it as v2 and start the trend again.
- 03Pick the three platformsGoogle AI Overviews, ChatGPT and Perplexity. Report each separately, because a blended number that hides a zero on one platform conceals the single most useful fact in the report.
- 04Set up clean sessionsLogged out or a clean profile, default model and settings, memory off, and a new chat for every run, never a reused thread. A logged-in session carries your own history and will flatter you. Set the client metro where the platform allows it, and record that you did.
- 05Run each question three times per platformThree runs, three fresh sessions. That is 108 observations. Spread them over several days rather than firing them all in one sitting, because the answers move.
- 06Count with the two-of-three ruleA question counts as won only if the business is named in at least two of its three runs. Named once out of three is variance rather than presence, and counting it is how honest measurement quietly turns into marketing.
- 07Record more than presencePer observation: named yes or no, the sentiment, which sources were cited, and which competitors appeared. The citation list is usually more actionable than your own score, because it tells you which surfaces the engine actually trusts in your market.
- 08Screenshot and date everythingAt minimum one run in five, plus every anomaly. An unscreenshotted claim about a generated answer is unfalsifiable, and three weeks later you will not be able to reproduce it.
- 09Report it as a fraction, per platform, with the method boxX/12 for each platform, the question-set version, run dates, total observations, and the words “directional measurement.” Then re-run monthly and read the trend, never a single month.
When a percentage becomes defensible
The quarterly instrument exists because the screen, by design, cannot answer “did the programme work.” It samples until the confidence interval is tight enough to make a claim: 10–15 money queries per topic-space, run enough times across the platform matrix to clear 400 observations: ten queries × ten runs × four surfaces, or fifteen × seven × four, or any arrangement that gets there, spread across at least five business days inside a two-week window.
That earns a percentage, reported as visibility-% ±5pp blended, with per-platform slices carried at their own weaker precision rather than borrowing the blended figure’s. And it earns a movement rule: a 5-point-or-greater change is movement we will claim, 2.5 to 5 points is a watch item reported but never attributed to anything, and under 2.5 points is noise we do not narrate at all.
Quarterly, not weekly, and that is a deliberate constraint rather than a scheduling convenience. When the underlying answers change every couple of days, a weekly reading tells you about the week. We would rather report four times a year and be right than twelve times and be entertaining.
What we will and will not report
Every run plan and every report we produce opens with this table, before a single number. It is the fastest way to tell an honest measurement from a decorative one, so ask whoever is quoting you to write theirs.
- Presence. How often you are named across a frozen question set, per platform, dated.
- Direction over time. Same questions, same method, repeated, read as a trend.
- Citation sources. Which domains the engines actually pull from in your market, and whether you control them.
- AI referral sessions and their conversion rate. Reported as counts and a rate, from your own analytics.
- Branded-search lift. The downstream fingerprint of being recommended.
- AI rank positions. They do not exist. See the research above.
- Prompt volumes. No platform publishes them; every vendor “volume” figure is an undisclosed model.
- Per-prompt positions over time. The citations churn faster than any reporting cycle.
- A complete AI-influenced customer journey. Nobody has this, including the people selling dashboards of it.
- Any ±1-point claim. The instrument does not have that resolution and we will not pretend it does.
A filled scorecard, so you know what you are getting
Illustration only. These are not a real client’s numbers. Your own figures come from the free Visibility Check, and our results for our own business are published on the tracking page.
How to audit anybody’s number, ours included
Ask four questions and the conversation resolves quickly. What were the twelve questions? If they will not show you the set, there may not be one. Which platforms, in what session state? A score run from a logged-in account with memory on is measuring the analyst. How many runs per question? One run is a coin flip; the research above is unambiguous about that. What does the number refuse to claim? A figure with no stated limits was not produced by anyone worried about being wrong.
Then apply it here. Our screen refuses to be a percentage at twelve questions. Our quarterly figure never appears without its confidence interval. Neither ever reports a rank. And we run the whole thing on ourselves and publish the result in whichever direction it goes. our own numbers are here, alongside what we refuse to promise, in writing, before anyone signs anything.
This page defines a metric and publishes its method. It does not sell anything. The free Visibility Check runs Tier 1 for your business at no cost. What it looks like when this measurement sits inside ongoing work is on the AI Visibility service page, and the category itself is defined at AI visibility.
Looking for what a “score” is and how to read one? That is the AI Visibility Score page. Short definitions live in the Visibility Glossary.
Fair questions about measuring this
Why avoid an AI rank tracker?
Because there is no rank to track. The SparkToro and Gumshoe research above found roughly a 1-in-1,000 chance of seeing the same brand list in the same order twice. A tool that reports your position is converting that randomness into a number, which feels like information and is not. What survives repeated sampling is presence, which is what we measure.
Isn't twelve questions a small sample?
Yes, and that is exactly why it is reported as a fraction and described as directional. Its job is triage, not proof: 1/12 and 8/12 point at completely different work. When a claim needs to carry weight, and has to answer whether the programme moved anything, we run the quarterly instrument at 400-plus observations and report the confidence interval with the number.
Why quarterly rather than weekly?
Because the thing being measured changes every couple of days. Ahrefs put AI Overview persistence at 2.15 days on average, with under 55% of cited URLs surviving from one observation to the next. Weekly readings of a surface that volatile produce narrative, not evidence, so we do not use them as reporting inputs.
Can I run this myself without buying anything?
Yes, and we would rather you did than took our word for it. Everything needed is on this page: the question-construction rule, the platform list, the session hygiene, the three runs, the two-of-three counting rule. It takes an afternoon. Doing it once makes you considerably harder to sell to, which we consider a fair trade.
What if my score goes down?
It sometimes will, including in months when the work was good, because these surfaces move for reasons nobody outside the model providers controls. That is exactly why the movement thresholds exist and why single months are not narrated. If a provider's numbers only ever go up, you are looking at a marketing artefact rather than a measurement.
Does being named actually bring in work?
Being named is presence, not revenue, and conflating the two would be the same dishonesty this page exists to avoid. What we can show alongside it is AI referral sessions and their conversion rate from your own analytics, plus branded-search movement. A rising X/12 with a flat phone is a real and useful finding, and it usually means the visibility is landing while the problem has moved somewhere else.
Get your Tier 1 screen, free
12 real buying questions for your trade and market, three assistants, three runs each: your X/12 per platform, the method box, the screenshots, and the first fix worth making. No sales call, and the findings are yours to keep.
Get my free Visibility Check