Home / Blog / We Track 6 AI Models Every Week
AI Visibility

We Track 6 AI Models Every Week. Here's Who They Recommend and Why

Storm Bennett · CEO, KillerSEOx · · 9 min read
Weekly AI share-of-voice table showing how often ChatGPT, Claude, Gemini, Perplexity, Grok, and Llama name four anonymized businesses: a pool retailer at 90 percent overall, a roofer at 61, a wildlife control company at 53, and a carpet cleaner at zero

Every week, our tracker asks ChatGPT, Claude, Gemini, Perplexity, Grok, and Meta Llama the same questions a real customer would ask. Questions like "who should I call to fix my roof in this town" and "recommend a company that removes moles from yards." Then it records who each model names, in what order, and what it says about them.

We run these checks across hundreds of tracked buyer prompts for the businesses on our platform, so at this point we have a pile of receipts most agencies don't: not opinions about how AI recommendations work, but the actual answers, week over week. This post is what that data looks like for four of those businesses, anonymized, and the patterns in why the models pick who they pick. Some of it matches what you'd guess. The best parts don't.

Key takeaways
  • AI visibility is wildly uneven: of four businesses on the same tracker, one gets named in 90% of checks, two sit near half, and one has never been named once in 90 model checks.
  • Perplexity is the toughest room. It posted the lowest mention rate of the six models for every single business we tracked: 60% where others hit 100, and two flat zeros.
  • When a model names a business at all, it named ours first about three times out of four, and inside the top three 94% of the time. AI answers barely have a page one. They have a podium.
  • The models say why they pick: they cite specific service pages, a BBB rating, review counts, years in business, and a street address. Every one of those is a signal you control.
  • Tiny wording changes flip results. One client is named by all six models on a specific buyer question and by zero of six on the generic "who's the best" version of the same question.

What did six AI models say when we asked who to hire?

Direct answer

Four businesses, same weekly tracker, four different stories: a pool and hot tub retailer was named in 90% of model checks, a coastal roofing contractor in 61%, a wildlife control company in 53%, and a carpet cleaning company in 0%. Ninety checks, not one mention. AI visibility is not evenly distributed.

Here's the latest weekly snapshot for four businesses we track, each in a different industry:

BusinessOverall visibilityRange across models
Pool & hot tub retailer, in business since the 1960s90%60% to 100%
Roofing contractor, New England coast61%0% to 100%
Wildlife control company, midwestern metro53%32% to 80%
Carpet cleaning company, midwestern metro0%0% everywhere

Two things in that table should stop you. First, the spread. These businesses are all real, all established, all doing work customers love. The difference between 90% and 0% isn't quality of service. It's how legible each business is to a machine trying to answer a buyer's question, which is the whole game I broke down in the AIO explainer.

Second, the range column. The roofer runs from 0% to 100% depending on which model you ask. A customer using ChatGPT hears about them every time. A customer using Perplexity never does. If you've only checked one assistant, you haven't checked your AI visibility. You've checked one sixth of it.

Which AI model is the hardest to get named by?

Direct answer

Perplexity, and it isn't close. It posted the lowest mention rate of the six models for every business we looked at: 60% on a client the other five models named nearly every time, 32% on another, and flat zeros on two more. It searches live, cites sources, and spreads its answers across more candidates.

Perplexity behaves differently because it is built differently. It runs a live web search on nearly every question, reads what comes back, and builds an answer with citations, usually naming several businesses instead of crowning one. More candidates per answer means a smaller share of voice for everybody. Our pool retailer, who the other five models named in essentially every check, gets named by Perplexity 60% of the time. The wildlife company drops to 32% there against 80% everywhere else.

I've come to treat Perplexity as the canary. Because it re-reads the live web every time, it reflects the current state of your pages, your citations, and your competitors' pages, not a memory of them. When Perplexity starts naming you consistently, it usually means the underlying signals got strong enough that the other models follow. The reverse is also useful: a business that's visible everywhere except Perplexity probably has a thin citation footprint and pages that don't survive being read by a machine in a hurry.

ChatGPT, meanwhile, is the swing voter. Across our four businesses it scored 100%, 100%, 40%, and 0%. When it knows you, it leads with you. When it doesn't, you simply don't exist in the answer.

Why do AI models recommend one business and skip another?

Direct answer

The models tell you. In their answers, they cite specific service pages, an A+ BBB rating, "50+ verified reviews," decades in business, a street address and phone number. They recommend businesses whose facts are easy to find, easy to extract, and easy to defend. Every one of those signals is publishable by you.

This is my favorite part of the data, because the models show their work. When they name our clients, the snippets read like a checklist of trust signals. Real excerpts from this cycle, lightly anonymized:

Look at what's actually being cited. Not the homepage. The models link the specific page that answers the specific question: the pool-closing service page when someone asks about pool closing, the town's service-area page when someone asks about roof repair in that town, the product category page when someone asks where to buy a robotic cleaner. The roofer gets named by five of six models for "roof repair in [their small town]" precisely because a page exists that says, in plain language, that they do that work in that place. This is the same mechanism that made service and town pages rank in Google, now paying a second dividend in AI answers.

The models aren't guessing. They're quoting. The businesses that get recommended are the ones that published something worth quoting.

And when a model names a business at all, it tends to commit. Across every mention in this cycle's data, our clients were named first about three times out of four, and inside the top three 94% of the time. There's no page two in an AI answer, and barely a middle. You're on the podium or you're absent.

Want to see this table for your business?

The free audit asks the six models about you and shows you what they say, alongside your Google rankings. About 60 seconds, no card, no signup.

Run my free audit

Why does the same business win one prompt and vanish on the next?

Direct answer

Because models answer the question actually asked. Our wildlife client is named by all six models, five of them first, for a specific buyer question naming the pest and the area. Ask the generic "who is the best [service] company in [city]" and zero of six name them. Specificity wins; superlatives scatter.

This is the pattern that changes how you should write pages. For the wildlife company, the question "best wildlife removal company near [city] for moles and groundhogs" gets them named by all six models, five at position one. The question "who is the best mole removal company in [city]" gets them named by none of the six. Nearly the same words. Opposite outcomes.

Why? The specific question has a specific answer, and this client owns it: dedicated pages for mole removal, for vole removal, for each service area, all extractable in one read. The generic "best company" question invites the model to hedge across directories, reviews, and lists, or to answer with criteria instead of names. Several models responded to the generic version with "here's how to choose a company" and named nobody at all.

The practical read: you can't own every phrasing, and you don't need to. You need to own the phrasings that carry buying intent, the ones that name a service, a problem, or a place. That's also why we track a fixed set of prompts weekly instead of spot-checking whatever comes to mind. Ask the same questions every week and movement means something. Ask different questions every time and you're reading tea leaves. If you want the ten-minute manual version, here's how to check what ChatGPT says about your business today.

What does zero AI visibility look like?

Direct answer

Ninety checks, zero mentions. Our carpet cleaning client was not named once by any of the six models across 15 tracked prompts. Nothing is broken and nobody is blacklisted. The models simply have no extractable reason to connect this business to those questions yet, so national brands soak up every answer.

The fourth row of that table deserves its own section, because more businesses live there than anywhere else. Fifteen prompts, six models, and not a single mention. If that business owner asked me "what do the AIs say about us," the honest answer is: nothing. Not something wrong. Nothing.

Here's what the zero actually teaches. The prompts this client tracks are mostly generic service questions without a strong local qualifier, things like "top rated carpet stretching services." Questions like that get answered with national franchises and how-to content, because when nothing anchors the question to a place, the model reaches for the biggest, most-cited entities it knows. A local company with a modest web footprint has no way into that answer. The fix isn't to shout louder into the same void. It's to compete where a local business can win: place-anchored questions, service-specific pages, and the citation and review footprint that gives a machine something to quote. Zero is not a verdict on the business. It's a to-do list.

I'd rather show you that zero than pretend every tracker screenshot is a highlight reel. Half the value of measuring is finding out where you actually stand before your competitor does.

How do you raise your AI share of voice?

Direct answer

Publish the facts the models quote: a page per service and per real service area, consistent name, address, and phone across the web, visible reviews, years in business, and schema that states it all machine-readably. Then measure weekly across all six models, because you can't manage a number you see once.

Everything the models cited in this data set is buildable. The work looks like this, roughly in order of leverage:

This is the exact loop the platform runs for the businesses in this post: fixed buyer prompts, six models, every week, with the misses turned into ranked actions. But the loop matters more than the tool. Whether you automate it or run it by hand on Friday afternoons, start measuring. The businesses in the 90% row aren't there because they're lucky. They're there because every question a buyer asks has a page, a fact, and a citation waiting to be quoted.

Quick answers

How do you measure AI share of voice?
Pick the questions a ready-to-buy customer would ask an AI assistant, ask all six major models those exact questions on a schedule, and record whether your business is named, where in the answer it appears, and who got named instead. The mention rate per model, tracked weekly, is your share of voice. The free audit runs the first check for you.
Do AI models give the same answer every time?
No. The same model can name you this week and skip you next week, and two models asked the identical question routinely disagree. That variance is why a single spot check tells you almost nothing. You need repeated checks over weeks to see your real mention rate instead of one lucky or unlucky answer.
Which AI model matters most for a local business?
ChatGPT has the most users, so start there. But buyers spread across all of them, and our data shows models disagree wildly about the same business. The one we watch closest is Perplexity: it posted the lowest mention rate for every business we tracked, so winning there usually means your citations and content are genuinely strong.
How long does it take to show up in AI answers?
When the underlying signals improve, movement can show up in web-search-backed models within weeks, because they retrieve live pages on every question. Models leaning on training data move slower. Fix the extractable facts first: service pages, consistent citations, reviews, and schema. Then track weekly so you actually see the change.
Storm Bennett, CEO of KillerSEOx
Storm Bennett is the CEO behind KillerSEOx. He's been getting businesses found since before Google sold ads.
← Back to all posts
Killerspots Agency