Skip to content
Mindflow Marketing — home
Call Youssef · (404) 775-9995 Free Visibility Check
Results Pricing
Get my free Visibility Check Call Youssef · (404) 775-9995
PUBLISHED IN FULL

The X/12 Protocol

This is the complete measurement protocol behind AI Share of Answer. It is published so that the number can be independently understood and reproduced: by you, by another agency, or by anyone who wants to check our work.

What is not published here: source code and automation credentials, client-confidential queries or data, internal QA checklists unrelated to the calculation, and vendor configuration that would introduce a security risk. None of those are needed to reproduce the number.

01

How the twelve questions are selected

The twelve are drawn from what a buyer actually types or asks when they are close to hiring, not from keyword volume. We build the candidate set from four sources: queries the client's own sales team reports hearing, the People Also Ask and related-question surfaces for the category, the phrasing competitors target in their own page titles, and the natural-language forms a buyer uses with an assistant, which are longer and more conversational than typed search.

From that candidate pool we select twelve that are commercial (the asker is choosing a provider, not researching a concept), answerable (a named business could legitimately appear in the response), and stable (the question will still be asked in six months). Twelve is a working compromise: enough for a percentage to mean something, few enough to run repeatedly by hand without sampling shortcuts.

02

How buyer intent and commercial relevance are validated

A question only enters the set if a reasonable answer to it could include a business like the client's. “How does a heat pump work” fails that test, because it is informational, and a good answer names no companies. “Who installs heat pumps in Marietta” passes.

We validate this by running each candidate once before the set is locked and reading what comes back. If the response names no providers at all, the question is measuring something other than commercial visibility and is replaced. This validation pass is disclosed in the baseline report along with any questions that were rejected and why.

03

Which surfaces are tested

Google AI Overviews, ChatGPT and Perplexity as the counted set, each scored and reported separately. A blended number that hides a zero on one platform conceals the single most useful fact in the report.

Google AI Mode is spot coverage only and is not part of the counted set. Copilot is not sampled. We report Bing indexation and Bing Places listing state instead, because presence there is verifiable while a sampled score would not be. A protocol that never names what it excludes is not a protocol.

04

Whether accounts, location and personalization are controlled

Yes, and this is where most published AI-visibility numbers quietly fall apart. Every run is executed signed out, in a fresh session with no conversation history, with browsing and memory features disabled where the platform allows it.

Location is set explicitly to the market being measured rather than inherited from the machine running the test, and the location used is recorded in the report. Where a platform does not permit location to be set deterministically, that limitation is stated against the affected surface rather than silently absorbed into the score.

05

When and how often tests run

The baseline runs before any work begins. After that, the full protocol runs quarterly, on the same twelve questions, under the same conditions.

Quarterly rather than monthly is deliberate. These surfaces move on their own between runs for reasons unrelated to anything an agency did, and a monthly cadence produces noise a client would reasonably mistake for progress. The monthly report carries the other two numbers; the X/12 figure updates on the quarter, and is dated.

06

How mentions, recommendations and citations differ

These are three different things and collapsing them is the most common way an AI-visibility number is inflated. We score them separately:

Mention. The business name appears in the response body. Recommendation. The response actively puts the business forward as an option for the asker. Citation. The business's own domain is linked or named as a source underpinning the answer.

A business can be cited without being recommended, and mentioned without either. The headline X/12 figure counts recommendations, because that is what a buying question is actually asking for. Mentions and citations are reported alongside as separate counts, never folded into the headline.

07

The exact X/12 calculation

X = the number of the twelve questions, on a given surface, where the business appears as a recommendation in the majority of runs. The denominator is always twelve. The result is reported per surface, never averaged across surfaces.

“Majority of runs” means at least two of three; see the next item. A business recommended in one run of three does not score; the appearance is recorded in the evidence file and noted as unstable, which is useful information but is not a point.

08

How inconsistent answers and repeated runs are handled

Each question is run three times per surface, in separate sessions. These systems are non-deterministic and a single run is close to meaningless.

A question scores when the business is recommended in at least two of the three runs. Where results split, recommended once and absent twice, the question is logged as unstable and the split is shown in the report rather than resolved by rounding. Instability is a finding in its own right: it usually means the business sits at the boundary of the model's consideration set, which is a different strategic position from being absent, and it is often the fastest thing to move.

09

How evidence is captured

Every run is captured as a full-text response with a timestamp, the surface, the location setting used, and the run number. Nothing is summarised at capture time.

The evidence file is delivered with the report and belongs to the client. It exists so the number can be checked rather than trusted: a client, or a client's other agency, can read the raw responses and verify the score was counted correctly. That is the point of publishing a protocol at all.

10

Known limitations and sources of variance

These systems change without notice. A model update can move results between runs for reasons that have nothing to do with any work performed. We do not claim otherwise, and a quarter where the number falls is reported as a number that fell, with what we know about why.

Twelve questions is a sample, not a census. It describes the buying questions selected, not every possible question in the category.

Personalization cannot be fully eliminated, only reduced. Signed-out, fresh-session, location-set conditions get close, and the residual variance is real.

Three runs is a small sample. It is enough to distinguish stable presence from noise; it is not enough to produce a confidence interval, and we do not present one.

Recommendation is a judgement call at the margin. Where a response mentions a business ambiguously, the call is made conservatively: ambiguous cases are scored as not-recommended, and flagged in the evidence file so the client can disagree.

Protocol version

X/12 Protocol v2.16 · last updated 26 July 2026. Changes to this protocol are versioned and dated. Where a change would break comparability with a client’s earlier baseline, the prior version is used for that client until the next baseline is re-run, and the report says which version produced the number.

THE METHOD

How we work and measure: the Entity Integration Framework.

Five stages, run in order, measured in public. In owner language: Measure, Fix, Build, Prove. This page shows the machinery underneath.

FIVE STAGES, IN ORDER

Identity → Architecture → Retrieval → Measurement → Outcome

01Identity

Engines can only recommend a business they can verify. Stage one makes you unambiguous: one exact name, address, and phone everywhere; a hardened Google Business Profile; entity signals and schema that agree with each other; every conflicting listing hunted down and fixed.

Who you are, provable everywhere.
02Architecture

Your website gets structured so both people and machines can use it: pages organized around how customers actually buy, answer-first formatting engines can quote, clean technical foundations, and crawler access open where it earns you visibility.

A site built to be read and cited.
03Retrieval

This is where you become the answer: citable content on the pages that matter, presence on the surfaces engines actually consult, reviews at a steady honest cadence, and the citation trail AI assistants follow when someone asks who to call.

Show up where the answer gets formed.
04Measurement

Everything is scored against your baseline, monthly: map-grid position city by city, the rankings that matter for your trade, and Share of Answer: the 12 real buying questions, scored X/12, never a percentage. The Work Ledger records every completed task.

Three numbers. One ledger. No mystery.
05Outcome

Visibility only counts when the phone rings. Stage five ties the numbers to calls and jobs, prunes what isn't earning, doubles down on what is, and sets the next quarter's priorities, in writing.

More of the right calls, proven.

In owner language: Measure. Fix. Build. Prove.

Measure is your baseline (stage 4 runs first and last). Fix is Identity and the urgent parts of Architecture. Build is Architecture and Retrieval. Prove is Measurement and Outcome: the Mindflow Monthly, three numbers, and the Work Ledger. Same machine, plain words.

Start with the free Visibility Check See public pricing
Call Youssef Free Visibility Check