Home / Blog / Weekly AI Visibility Report
Product

What an AI Visibility Tool Should Show You Every Week

Storm Bennett · CEO, KillerSEOx · · 9 min read
A weekly AI visibility report for a pool retailer showing per-model mention rates across six AI models, ChatGPT and Grok at 47 percent down to Claude at 33 percent, with a 39 percent overall score and a prompt-level split of five questions won and twelve at zero

Plenty of tools will now hand you an AI visibility score. One number, maybe a trend arrow, and an invoice. Here is the problem: a single score across every AI model and every question is exactly as useful as a single average position across every keyword, which is to say almost not at all. We run a weekly six-model check on every business we track, and the value has never once been the score. It is the five layers underneath it.

Every week our platform asks ChatGPT, Claude, Gemini, Perplexity, Grok, and Meta's Llama the questions a real customer would ask, then records who each model names, in what order, and in what words. For one pool retailer we track, that is 30 prompts across 6 models, 180 recorded answers, every single week. This post walks through what that report actually contains, using the real numbers from a recent week, because the difference between a vanity score and a work order is everything below the headline.

Key takeaways
  • A useful AI visibility report has five layers: per-model mention rates, per-prompt results, your position inside each answer, the actual quoted snippet, and the week-over-week change. A single blended score hides all five.
  • Models disagree with each other more than most people expect. In the same week, the same roofer was mentioned in 50% of ChatGPT answers and 14% of Perplexity answers. A one-model check would have gotten this business completely wrong in either direction.
  • The sharpest pattern in every report we run: specific buying questions get you named, broad generic questions get you skipped. Our pool retailer went 100% across all six models on five specific prompts and 0% on twelve generic ones in the same run.
  • The snippet layer tells you why you were mentioned: the sources the model leaned on, a dealer page, a BBB profile, a service page with the city in it. That is the layer that turns a score into a to-do list.
  • A zero is not a verdict, it is a starting line. One client measured 0% across 90 checks. That does not mean AI hates them. It means the models have nothing about them worth citing, which is a fixable content and citations problem.

What is a weekly AI visibility report?

Direct answer

A weekly AI visibility report measures how often AI assistants name your business when asked the questions your customers ask. A real one tracks the same set of prompts across multiple models every week and records four things per answer: whether you were mentioned, your position in the answer, the exact wording, and the sources cited.

The prompts are the whole game. We do not track "tell me about [business name]," because your own brand name is a question you have already won. We track the questions someone with money asks before they know you exist: "Who offers pool closing service in Cincinnati?" "Can you help me find a roofing contractor in Hampton Falls?" "Best wildlife removal company near Cincinnati for moles and groundhogs?" Those are real prompts from our fleet, and they are the AI equivalent of money keywords. I covered how assistants build these answers in the AIO pillar post, and how to see your own starting point in what ChatGPT says about your business.

Each answer gets scored, not just counted. Being the first business named in a six-option list is a different outcome than being the fifth, the same way position 1 and position 5 are different outcomes on a results page. And the report keeps the receipt: the actual sentence the model wrote, with the links it cited. When a model says a business is "A+ BBB rated with over 50 verified reviews," you know exactly which source earned that mention, because the model just told you.

Why check six models instead of one?

Direct answer

Because the models genuinely disagree. In one week, the same roofer showed up in 50% of ChatGPT answers and 14% of Perplexity answers to identical prompts. Each model searches differently and trusts different sources, so a one-model check gives you one model's opinion, not your AI visibility.

People assume the big models roughly agree. Our data says otherwise, week after week. A roofer we track answered 22 prompts on each of 6 models in one recent run. ChatGPT named them in 50% of answers. Grok said 41%. Claude 27%, Gemini and Llama 23% each, and Perplexity just 14%. Same business, same week, same questions, and the spread between the friendliest model and the harshest one was 36 points.

The pool retailer's spread ran tighter but still real: ChatGPT and Grok at 47%, Gemini, Perplexity, and Llama at 37%, Claude at 33%, blending to a 39% overall score. If that retailer had checked only Claude, they would think they were losing. Only ChatGPT, and they would think they were fine. Neither number is the truth. The truth is the spread itself, because your next customer might be asking any of the six, and who each model recommends and why turns out to be its own study.

The spread is also diagnostic. Perplexity leans hardest on live web search and citations, so when Perplexity is your weakest model, thin citations and weak directory presence are usually the reason. Models that rely more on trained knowledge reward businesses that have been written about consistently for years. Where you are weak tells you what kind of evidence you are missing, and you cannot see any of that in a blended score.

What should the per-prompt breakdown tell you?

Direct answer

Which questions you win, which you lose, and the pattern that separates them. Our pool retailer hit 100% across all six models on five specific prompts and 0% on twelve broad ones in the same week. Specific, evidenced questions get businesses named. Generic questions get generic advice.

This is the layer where the report stops being a scoreboard and starts being a strategy. In that pool retailer's week, five prompts came back mentioned by all six models, most at position 1: the pool closing question, the pool supplies question, the best pool store question, a brand-of-hot-tub dealer question, and where to buy a swimming pool. On the dealer prompt, every single model named them first, because the manufacturer's own dealer page says so and every model found it.

Now the other half. Twelve of the thirty prompts came back 0% across all six models. And here is the part that should change how you think about AI visibility: the losing prompts were the biggest ones. "Where can I buy pool supplies like chemicals, filters, and testing kits from a reliable pool store?" is the highest-volume prompt on their board, and it went 0 for 6. "What's the best swimming pool company near me for design, installation, and service?" also 0 for 6. The roofer shows the same cliff: 100% on "find a roofing contractor in Hampton Falls," with all six models naming them first, and 0% on "how much does a new roof cost in New Hampshire," the single highest-volume question a roofing customer asks.

Specific question, you get named. Generic question, you get skipped. That split shows up in every report we run, and it is the most actionable pattern in AI search.

Why it happens is no mystery once you read the snippets. On specific prompts, the models are quoting evidence: a service page titled for the exact job and city, a dealer listing, a BBB profile, a review count. On generic prompts there is no evidence that says any particular business is the answer, so the model either lists whoever has the strongest general footprint or answers with advice and names nobody. The fix is not "do AI optimization" in the abstract. It is: build the page that answers the generic question with checkable facts, and make sure the sources models trust say what you need them to say. That is exactly the work in the AIO on-page checklist, and it is why the cost question the roofer keeps losing is a content assignment, not a mystery.

Position matters inside the wins too. On the wildlife control client we track, the best prompt came back 100% with position 1 on nearly every model. Being named first in an AI answer is the new top of page one, and the difference between first and fifth in a recommendation list decides who gets the call. The parallel with how AI Overviews and the local pack split the click is direct: presence is table stakes, position is the prize.

Want to see your own six-model report?

The free audit includes an AI visibility check alongside the 21-point SEO scan, in about 60 seconds, no card required. See which buying questions you win and which ones name your competitor instead.

Run a free audit

What does a zero actually mean?

Direct answer

A zero means the models found nothing about you worth citing, not that they rejected you. One client we track measured 0% across 90 checks in a week. Their fixable causes were thin citations, no dedicated service pages, and a weak review footprint. A zero is a baseline, and baselines move.

One client in our fleet, a carpet cleaner, ran 15 prompts across all 6 models in a week: 90 answers, 0 mentions. Not one model named them once. If you got that as a bare score, it would feel like a verdict. Read as a report, it is a diagnosis. The models were not avoiding this business. They had never heard of it in any source they trust: sparse citations, no service pages targeting the questions being asked, and competitors with years of accumulated directory and review presence soaking up every mention.

That list of causes is a work order, and none of it is exotic. A citation audit tells you what the directories say about you today. Service pages built for the actual questions give models something to quote. Reviews and profiles give them the trust signals they cite by name, and we watch models literally write "A+ BBB rating" and "50+ verified reviews" into their answers for the clients who have them. The zero is where you start. The weekly cadence is how you catch the first mention when it lands, and on a new or invisible business, that first mention tends to show up on one model weeks before the others follow.

What should you do with the report each week?

Direct answer

Ten minutes, four questions: did the overall trend move, which model changed most, which prompts flipped, and what did the snippets cite. Then feed the losing prompts into your content queue and the missing sources into your citations work. The report is an input to next week's work, not a trophy.

Here is the loop we run on every account, and the reason the report exists at all:

  1. Read the trend, not the score. 39% means little on its own. Up from 31% three weeks running means your work is landing. AI answers move faster than Google rankings, which is exactly why the check is weekly.
  2. Find the flipped prompts. A prompt that went from mentioned to absent on multiple models is your early warning: a competitor got written about, a source changed, an answer got rebuilt. You want that signal the week it happens, not at quarter end.
  3. Read the snippets on wins you did not expect. The model tells you which source earned the mention. Whatever it cited, strengthen it. That dealer page carrying five perfect scores? It is doing more for that retailer in AI search than most of their website.
  4. Turn the losing prompts into content assignments. The roofer losing the cost question does not need a better score, they need a cost page a model can quote. Every 0% prompt with real volume is a title for the content queue.
  5. Route the source gaps to citations and reviews. When your weakest model is the one that leans on live search, the fix lives in directories and profiles, not on your website.

None of this requires believing any hype about AI replacing search. It requires noticing that some real slice of your customers now asks an assistant instead of scrolling results, and that who gets named in those answers is measurable, weekly, for every model that matters. We built exactly this report into the platform, and the free tier will run it on your business without a card. The score is the headline. The five layers under it are the reason to look.

Quick answers

What is a good AI visibility score?
There is no universal good score, because the number depends entirely on which prompts you track. A business tracking only its own brand name will score near 100 and learn nothing. Across the local businesses we track weekly, overall scores run from 0 to the mid 50s, and the useful signal is the split underneath: which buying questions you win, which you lose, and whether the trend is up.
How often should you check your AI visibility?
Weekly. AI answers are not stable: models get updated, their search indexes refresh, and the sources they cite change constantly. A monthly check cannot tell you whether a fix worked or when a competitor displaced you. A weekly run across the same prompts gives you a real trend line while staying cheap enough to sustain.
Why do AI models mention my competitor but not me?
Because the models found better evidence for the competitor. Assistants that search the web build answers from sources they can read and trust: service pages that name the service and the city, directory and BBB listings, review profiles, manufacturer dealer pages. If your competitor has a dedicated page for the exact question and you have a thin homepage, they get cited and you do not.
How do I check what AI models say about my business for free?
You can ask each assistant buying questions by hand and note who gets named, which works but does not scale past a few prompts. Our free audit includes an AI visibility check alongside the SEO scan, and the free tier of the platform runs ongoing checks without a credit card. Either way, test the questions a customer would ask, not your brand name.
Storm Bennett, CEO of KillerSEOx
Storm Bennett is the CEO behind KillerSEOx. He's been getting businesses found since before Google sold ads.
← Back to all posts
Killerspots Agency