Open protocol · v1.0

The AI Visibility Benchmark

An open, reproducible method for measuring whether AI assistants name your business — and whether they get it right.

By Found ClearlyPublished August 13, 20269 min read

Abstract

The AI Visibility Benchmark measures how often and how accurately AI assistants name a local service business in response to real buyer questions. A run is 15 prompts across five buyer intents, put to ChatGPT, Gemini, and Perplexity in clean sessions, and scored 0 to 4 per answer — where a mention carrying a wrong fact scores below a clean mention rather than above an absence. The output is a per-assistant Visibility Score from 0 to 100, plus the list of competitors the assistants named instead. The method is published in full so any business can run it, repeat it monthly, and compare its own results over time.

Key points

  • A single query to one assistant is an anecdote. The protocol requires 15 prompts, three assistants, and three runs across three separate days before a number is reported.
  • Being named with a wrong service area or a service you do not sell scores 1 out of 4 — below a clean mention, because the buyer acts on the error.
  • Scores are reported per assistant and never averaged into one figure, because merging them hides the assistant you are absent from.
  • The most useful output is usually not your own score but the competitor list: the businesses the assistants treat as your peer set.
  • The benchmark measures visibility, not revenue, and running it does not cause an assistant to cite anyone.

Ask a marketing supplier whether your business is visible in AI search and you will get an answer. Ask how they measured it and the answer usually thins out: someone typed a question into ChatGPT once, saw the business name, and reported a win — or did not see it, and reported a crisis. Both readings are worthless, and for the same reason. One sample from a non-deterministic system tells you nothing you can act on or repeat.

So we wrote down the method we use, in enough detail that anyone can run it without us. It costs nothing, needs no API keys, and works on any local service business. You can point it at your own company, at a competitor, or at us.

What the benchmark actually measures

It answers two separate questions that are usually collapsed into one. First: when a buyer who has never heard of you asks an assistant who does this work, are you named? Second: when someone who has heard of you asks what you do, is the answer correct? A business can pass the first and fail the second, and that combination is worse than being invisible — the buyer receives a confident description of a service area you do not cover and never calls to check.

Separating presence from accuracy is the single most important design decision in the rubric. A mention that misstates your coverage, your services, or your phone number scores below a clean mention, not above an absence.

The prompt set: five buyer intents, three phrasings

Fifteen prompts is a compromise, chosen on purpose. Fewer and one lucky answer carries the score. More and nobody runs it a second time. The five classes cover the moments where an assistant can decide for or against you: unbranded discovery, shortlist qualification, attribute filtering, branded fact-checking, and head-to-head comparison. Each gets three phrasings because real buyers do not ask the same question twice, and assistants are sensitive to wording in ways that a single template would hide.

The prompts are questions, not keyword strings. Nobody types 'best HVAC Fort Worth' into an assistant; they type a sentence describing their problem. A benchmark built on keyword fragments measures something no customer is doing.

Session hygiene is most of the validity

The commonest way to get a flattering, useless result is to run the test from your own account. Memory, chat history, custom instructions, and location personalisation all push the assistant toward naming the business you have been discussing for weeks. The rules below exist so that the run measures the business rather than the tester.

State the location in the prompt rather than relying on device location, and the run becomes reproducible from any office in the world — which is what makes a second person's result comparable to yours.

Scoring, and what to do with the number

Score each answer 0 to 4, total the run, and express it as a percentage of the maximum (15 prompts × 4). Report it per assistant. Averaging ChatGPT, Gemini, and Perplexity into one figure feels tidier and destroys the only actionable part of the result: they are different markets with different source preferences, and the interesting finding is almost always that you are present in one and absent from another.

Then read the competitor column. For most businesses the score is a blunt trend line, while the named-competitor list is a direct instruction: these are the companies the assistants currently treat as the answer to your buyer's question. That tells you which sites to study and which corroborating sources the engines are leaning on.

How often to run it

  1. Run the full set three times across three separate days. Record every answer verbatim with its date and the model version shown in the interface.
  2. Take the median of the three runs as the reported score, not the best one.
  3. Repeat monthly. A quarterly cadence cannot separate your own changes from a model update.
  4. Log every material change you make to the site alongside the run dates, so a movement has candidate causes.
  5. Re-read the competitor list each month. It moves faster than the score.

What a low score does and does not tell you

A low unbranded score usually means the assistants have too little corroborated material about the business to risk naming it — a thin site, few independent references, or content that answers no question anyone asked. A low branded-accuracy score usually means the opposite problem: plenty of material, but stale, contradictory, or spread across profiles that disagree with each other.

Neither reading is a promise about what happens after you fix it. Answer engines choose their own sources, and no method — this one included — makes a citation happen. What a benchmark buys you is the ability to tell whether anything you did made a difference, which is more than most marketing reporting can claim.

Protocol · Step 1

The 15-prompt set

Five buyer intents, three phrasings each. Fill in your own details to generate the run — everything happens in your browser.

Nothing here is sent anywhere — the prompts and the sheet are generated in your browser. Run each prompt in a fresh session on ChatGPT, Google Gemini, Perplexity, then score every answer 0–4.

  1. discovery-1 · Unbranded discovery · unbranded

    Who are the best emergency AC repair companies in Fort Worth?

  2. discovery-2 · Unbranded discovery · unbranded

    I need emergency AC repair in Fort Worth. Which local companies should I look at?

  3. discovery-3 · Unbranded discovery · unbranded

    Recommend three reliable emergency AC repair providers serving Fort Worth.

  4. qualification-1 · Shortlist qualification · unbranded

    I need emergency AC repair in Fort Worth this week. Give me three companies and one reason for each.

  5. qualification-2 · Shortlist qualification · unbranded

    Which emergency AC repair company in Fort Worth would you call first, and why?

  6. qualification-3 · Shortlist qualification · unbranded

    Shortlist two emergency AC repair providers in Fort Worth for a homeowner who cares about reliability over price.

  7. attribute-1 · Attribute filtering · unbranded

    Which emergency AC repair companies in Fort Worth offer same-day appointments?

  8. attribute-2 · Attribute filtering · unbranded

    Who does emergency AC repair in Fort Worth and is transparent about pricing?

  9. attribute-3 · Attribute filtering · unbranded

    Which emergency AC repair providers near Fort Worth are licensed and well reviewed?

  10. branded-fact-1 · Branded fact check · branded

    What does Your Business Name do, and where do they operate?

  11. branded-fact-2 · Branded fact check · branded

    Is Your Business Name a good choice for emergency AC repair in Fort Worth?

  12. branded-fact-3 · Branded fact check · branded

    What services does Your Business Name offer, and how do I contact them?

  13. comparison-1 · Head-to-head comparison · branded

    Compare Your Business Name with other emergency AC repair companies in Fort Worth.

  14. comparison-2 · Head-to-head comparison · branded

    What are the alternatives to Your Business Name for emergency AC repair in Fort Worth?

  15. comparison-3 · Head-to-head comparison · branded

    Why might someone choose Your Business Name over another emergency AC repair provider in Fort Worth?

discoveryUnbranded

Unbranded discovery

A buyer who does not know you exists asks who does this work.

Why it earns a slot: This is the class that decides whether you are in the market at all. Everything else measures how you are handled once you have been found.

qualificationUnbranded

Shortlist qualification

A buyer with urgency asks for a small, justified shortlist rather than a list.

Why it earns a slot: Assistants behave differently when asked to choose and justify. Appearing in a long list but never in a shortlist is a distinct, diagnosable failure.

attributeUnbranded

Attribute filtering

A buyer with a specific constraint — cost, speed, licensing, availability.

Why it earns a slot: Constraint questions expose whether the assistant has any substantive facts about you beyond a name, which is where thin content shows up.

branded-factBranded

Branded fact check

Someone who has heard your name checks what you actually do.

Why it earns a slot: Measures accuracy rather than presence. A confident wrong answer about your service area is worse than silence, and only this class catches it.

comparisonBranded

Head-to-head comparison

A buyer comparing you against the alternatives they have found.

Why it earns a slot: Reveals which competitors the assistant treats as your peer set — often the most actionable output of the whole run.

Protocol · Step 2

The scoring rubric

Score every answer once, on this five-point scale. The gap between level 1 and level 2 is the point of the whole rubric: a mention carrying a wrong fact costs you the customer, so it must not outscore a clean one.

ScoreLevelWhat it means
0AbsentThe business is not named anywhere in the answer.
1Named with an errorThe business is named, but a material fact is wrong — service area, services offered, contact route, or status.
2Named correctlyThe business is named and everything stated about it is accurate, but it is listed among others without distinction.
3Named with a reasonThe business is named accurately and the answer gives a specific reason to consider it.
4Named first, with a linkThe business leads the answer and the assistant links or cites a page on its own domain.

A full run is 15 prompts against one assistant, so the maximum is 60 points. Report the total as a percentage of that maximum, per assistant.

Protocol · Step 3

Session rules

Most of a benchmark's validity lives here. Run the test from your own everyday account and you will measure your account history, not your business.

  1. 1Start a new conversation for every prompt. Answers within one thread contaminate the next.
  2. 2Disable memory, chat history training, and custom instructions on any account that offers them.
  3. 3Do not sign in where the assistant works signed out, and never run from an account that has previously discussed the business.
  4. 4State the location in the prompt rather than relying on device location, so a run from any office is reproducible.
  5. 5Never hint the answer. Do not paste the business name into an unbranded prompt, even as a follow-up.
  6. 6Record the answer verbatim, with the date, the assistant, and the model version shown in the interface.
  7. 7Run the full set three times across three separate days before reporting a number.

ChatGPT

Run with web search enabled, in a session with memory and custom instructions turned off.

Google Gemini

Run signed out where possible; a signed-in session applies personalisation you cannot reproduce.

Perplexity

Use the default model and record the cited sources alongside the score.

What this cannot tell you

  • Assistant answers are non-deterministic — the same prompt can return a different answer minutes later. Three runs across three days is the minimum that makes a reported number meaningful.
  • Personalisation, region, and account history shift results, so a score is comparable to your own previous scores rather than to another company's published figure.
  • Model versions change without notice. A month-on-month drop can be a model update rather than anything the business did.
  • The benchmark measures whether a business is named, not whether a customer called. It is a visibility measurement, not a revenue measurement.
  • Nothing in the method causes an assistant to cite a business. It establishes where you stand so the work can be aimed at something real.

How to cite this

Found Clearly. "The AI Visibility Benchmark", version 1.0. Published 2026-08-13.

Free to run, quote, and adapt with attribution and a link to this page — including by competitors, and including against us.

Primary sources

FAQ

Related questions

How do I measure whether ChatGPT recommends my business?

Run a fixed set of buyer questions in fresh sessions with memory and custom instructions disabled, score each answer from 0 to 4 for whether the business is named and whether the facts are correct, and repeat the set three times across three days. A single query is not a measurement, because assistant answers vary between runs.

Why does the protocol use 15 prompts rather than one or two?

Assistant answers are non-deterministic and sensitive to phrasing. Fifteen prompts across five buyer intents means one unusually good or bad answer cannot dominate the result, while keeping the run short enough that a business owner will repeat it monthly.

Should I average my scores across ChatGPT, Gemini, and Perplexity?

No. Report them separately. The assistants draw on different sources, and the most useful finding is usually that a business is present in one and absent from another — which a single averaged number hides.

Can running this benchmark improve my AI visibility?

Measuring changes nothing on its own. The benchmark tells you where you currently stand and which competitors the assistants name instead, so that any work you do afterwards can be checked against a baseline rather than assumed to have worked.

Is the benchmark free to use commercially?

Yes. Run it for your own business or for clients, quote the rubric, and adapt the prompt set, with attribution and a link back to this page.