Ask a marketing supplier whether your business is visible in AI search and you will get an answer. Ask how they measured it and the answer usually thins out: someone typed a question into ChatGPT once, saw the business name, and reported a win — or did not see it, and reported a crisis. Both readings are worthless, and for the same reason. One sample from a non-deterministic system tells you nothing you can act on or repeat.
So we wrote down the method we use, in enough detail that anyone can run it without us. It costs nothing, needs no API keys, and works on any local service business. You can point it at your own company, at a competitor, or at us.
What the benchmark actually measures
It answers two separate questions that are usually collapsed into one. First: when a buyer who has never heard of you asks an assistant who does this work, are you named? Second: when someone who has heard of you asks what you do, is the answer correct? A business can pass the first and fail the second, and that combination is worse than being invisible — the buyer receives a confident description of a service area you do not cover and never calls to check.
Separating presence from accuracy is the single most important design decision in the rubric. A mention that misstates your coverage, your services, or your phone number scores below a clean mention, not above an absence.
The prompt set: five buyer intents, three phrasings
Fifteen prompts is a compromise, chosen on purpose. Fewer and one lucky answer carries the score. More and nobody runs it a second time. The five classes cover the moments where an assistant can decide for or against you: unbranded discovery, shortlist qualification, attribute filtering, branded fact-checking, and head-to-head comparison. Each gets three phrasings because real buyers do not ask the same question twice, and assistants are sensitive to wording in ways that a single template would hide.
The prompts are questions, not keyword strings. Nobody types 'best HVAC Fort Worth' into an assistant; they type a sentence describing their problem. A benchmark built on keyword fragments measures something no customer is doing.
Session hygiene is most of the validity
The commonest way to get a flattering, useless result is to run the test from your own account. Memory, chat history, custom instructions, and location personalisation all push the assistant toward naming the business you have been discussing for weeks. The rules below exist so that the run measures the business rather than the tester.
State the location in the prompt rather than relying on device location, and the run becomes reproducible from any office in the world — which is what makes a second person's result comparable to yours.
Scoring, and what to do with the number
Score each answer 0 to 4, total the run, and express it as a percentage of the maximum (15 prompts × 4). Report it per assistant. Averaging ChatGPT, Gemini, and Perplexity into one figure feels tidier and destroys the only actionable part of the result: they are different markets with different source preferences, and the interesting finding is almost always that you are present in one and absent from another.
Then read the competitor column. For most businesses the score is a blunt trend line, while the named-competitor list is a direct instruction: these are the companies the assistants currently treat as the answer to your buyer's question. That tells you which sites to study and which corroborating sources the engines are leaning on.
How often to run it
- Run the full set three times across three separate days. Record every answer verbatim with its date and the model version shown in the interface.
- Take the median of the three runs as the reported score, not the best one.
- Repeat monthly. A quarterly cadence cannot separate your own changes from a model update.
- Log every material change you make to the site alongside the run dates, so a movement has candidate causes.
- Re-read the competitor list each month. It moves faster than the score.
What a low score does and does not tell you
A low unbranded score usually means the assistants have too little corroborated material about the business to risk naming it — a thin site, few independent references, or content that answers no question anyone asked. A low branded-accuracy score usually means the opposite problem: plenty of material, but stale, contradictory, or spread across profiles that disagree with each other.
Neither reading is a promise about what happens after you fix it. Answer engines choose their own sources, and no method — this one included — makes a citation happen. What a benchmark buys you is the ability to tell whether anything you did made a difference, which is more than most marketing reporting can claim.