We Track 6 AI Models Every Week. Here's Who They Recommend and Why
Every week, our tracker asks ChatGPT, Claude, Gemini, Perplexity, Grok, and Meta Llama the same questions a real customer would ask. Questions like "who should I call to fix my roof in this town" and "recommend a company that removes moles from yards." Then it records who each model names, in what order, and what it says about them.
We run these checks across hundreds of tracked buyer prompts for the businesses on our platform, so at this point we have a pile of receipts most agencies don't: not opinions about how AI recommendations work, but the actual answers, week over week. This post is what that data looks like for four of those businesses, anonymized, and the patterns in why the models pick who they pick. Some of it matches what you'd guess. The best parts don't.
- AI visibility is wildly uneven: of four businesses on the same tracker, one gets named in 90% of checks, two sit near half, and one has never been named once in 90 model checks.
- Perplexity is the toughest room. It posted the lowest mention rate of the six models for every single business we tracked: 60% where others hit 100, and two flat zeros.
- When a model names a business at all, it named ours first about three times out of four, and inside the top three 94% of the time. AI answers barely have a page one. They have a podium.
- The models say why they pick: they cite specific service pages, a BBB rating, review counts, years in business, and a street address. Every one of those is a signal you control.
- Tiny wording changes flip results. One client is named by all six models on a specific buyer question and by zero of six on the generic "who's the best" version of the same question.
What did six AI models say when we asked who to hire?
Four businesses, same weekly tracker, four different stories: a pool and hot tub retailer was named in 90% of model checks, a coastal roofing contractor in 61%, a wildlife control company in 53%, and a carpet cleaning company in 0%. Ninety checks, not one mention. AI visibility is not evenly distributed.
Here's the latest weekly snapshot for four businesses we track, each in a different industry:
| Business | Overall visibility | Range across models |
|---|---|---|
| Pool & hot tub retailer, in business since the 1960s | 90% | 60% to 100% |
| Roofing contractor, New England coast | 61% | 0% to 100% |
| Wildlife control company, midwestern metro | 53% | 32% to 80% |
| Carpet cleaning company, midwestern metro | 0% | 0% everywhere |
Two things in that table should stop you. First, the spread. These businesses are all real, all established, all doing work customers love. The difference between 90% and 0% isn't quality of service. It's how legible each business is to a machine trying to answer a buyer's question, which is the whole game I broke down in the AIO explainer.
Second, the range column. The roofer runs from 0% to 100% depending on which model you ask. A customer using ChatGPT hears about them every time. A customer using Perplexity never does. If you've only checked one assistant, you haven't checked your AI visibility. You've checked one sixth of it.
Which AI model is the hardest to get named by?
Perplexity, and it isn't close. It posted the lowest mention rate of the six models for every business we looked at: 60% on a client the other five models named nearly every time, 32% on another, and flat zeros on two more. It searches live, cites sources, and spreads its answers across more candidates.
Perplexity behaves differently because it is built differently. It runs a live web search on nearly every question, reads what comes back, and builds an answer with citations, usually naming several businesses instead of crowning one. More candidates per answer means a smaller share of voice for everybody. Our pool retailer, who the other five models named in essentially every check, gets named by Perplexity 60% of the time. The wildlife company drops to 32% there against 80% everywhere else.
I've come to treat Perplexity as the canary. Because it re-reads the live web every time, it reflects the current state of your pages, your citations, and your competitors' pages, not a memory of them. When Perplexity starts naming you consistently, it usually means the underlying signals got strong enough that the other models follow. The reverse is also useful: a business that's visible everywhere except Perplexity probably has a thin citation footprint and pages that don't survive being read by a machine in a hurry.
ChatGPT, meanwhile, is the swing voter. Across our four businesses it scored 100%, 100%, 40%, and 0%. When it knows you, it leads with you. When it doesn't, you simply don't exist in the answer.
Why do AI models recommend one business and skip another?
The models tell you. In their answers, they cite specific service pages, an A+ BBB rating, "50+ verified reviews," decades in business, a street address and phone number. They recommend businesses whose facts are easy to find, easy to extract, and easy to defend. Every one of those signals is publishable by you.
This is my favorite part of the data, because the models show their work. When they name our clients, the snippets read like a checklist of trust signals. Real excerpts from this cycle, lightly anonymized:
- "With an A+ Better Business Bureau rating and over 50 verified customer reviews..."
- "Serving the area since 1966..." and, for another client, "founded in the late 1990s, 25+ years in business..."
- "You can visit their showroom at [street address] or call [phone number]..."
- "They have a dedicated section for robotic cleaners with many in-stock models..."
Look at what's actually being cited. Not the homepage. The models link the specific page that answers the specific question: the pool-closing service page when someone asks about pool closing, the town's service-area page when someone asks about roof repair in that town, the product category page when someone asks where to buy a robotic cleaner. The roofer gets named by five of six models for "roof repair in [their small town]" precisely because a page exists that says, in plain language, that they do that work in that place. This is the same mechanism that made service and town pages rank in Google, now paying a second dividend in AI answers.
And when a model names a business at all, it tends to commit. Across every mention in this cycle's data, our clients were named first about three times out of four, and inside the top three 94% of the time. There's no page two in an AI answer, and barely a middle. You're on the podium or you're absent.
Want to see this table for your business?
The free audit asks the six models about you and shows you what they say, alongside your Google rankings. About 60 seconds, no card, no signup.
Run my free auditWhy does the same business win one prompt and vanish on the next?
Because models answer the question actually asked. Our wildlife client is named by all six models, five of them first, for a specific buyer question naming the pest and the area. Ask the generic "who is the best [service] company in [city]" and zero of six name them. Specificity wins; superlatives scatter.
This is the pattern that changes how you should write pages. For the wildlife company, the question "best wildlife removal company near [city] for moles and groundhogs" gets them named by all six models, five at position one. The question "who is the best mole removal company in [city]" gets them named by none of the six. Nearly the same words. Opposite outcomes.
Why? The specific question has a specific answer, and this client owns it: dedicated pages for mole removal, for vole removal, for each service area, all extractable in one read. The generic "best company" question invites the model to hedge across directories, reviews, and lists, or to answer with criteria instead of names. Several models responded to the generic version with "here's how to choose a company" and named nobody at all.
The practical read: you can't own every phrasing, and you don't need to. You need to own the phrasings that carry buying intent, the ones that name a service, a problem, or a place. That's also why we track a fixed set of prompts weekly instead of spot-checking whatever comes to mind. Ask the same questions every week and movement means something. Ask different questions every time and you're reading tea leaves. If you want the ten-minute manual version, here's how to check what ChatGPT says about your business today.
What does zero AI visibility look like?
Ninety checks, zero mentions. Our carpet cleaning client was not named once by any of the six models across 15 tracked prompts. Nothing is broken and nobody is blacklisted. The models simply have no extractable reason to connect this business to those questions yet, so national brands soak up every answer.
The fourth row of that table deserves its own section, because more businesses live there than anywhere else. Fifteen prompts, six models, and not a single mention. If that business owner asked me "what do the AIs say about us," the honest answer is: nothing. Not something wrong. Nothing.
Here's what the zero actually teaches. The prompts this client tracks are mostly generic service questions without a strong local qualifier, things like "top rated carpet stretching services." Questions like that get answered with national franchises and how-to content, because when nothing anchors the question to a place, the model reaches for the biggest, most-cited entities it knows. A local company with a modest web footprint has no way into that answer. The fix isn't to shout louder into the same void. It's to compete where a local business can win: place-anchored questions, service-specific pages, and the citation and review footprint that gives a machine something to quote. Zero is not a verdict on the business. It's a to-do list.
I'd rather show you that zero than pretend every tracker screenshot is a highlight reel. Half the value of measuring is finding out where you actually stand before your competitor does.
How do you raise your AI share of voice?
Publish the facts the models quote: a page per service and per real service area, consistent name, address, and phone across the web, visible reviews, years in business, and schema that states it all machine-readably. Then measure weekly across all six models, because you can't manage a number you see once.
Everything the models cited in this data set is buildable. The work looks like this, roughly in order of leverage:
- A real page for each service and each real service area. The specific page is what gets cited. The roofer's small-town pages and the retailer's service pages did the heavy lifting in this cycle.
- Trust facts stated in plain text. Years in business, ratings, review counts, address, phone. The models quoted every one of these verbatim. If those facts live only in a footer image or a PDF, they don't exist.
- The on-page machinery. Entity schema, extractable answers, llms.txt, crawler access. That's a checklist of its own, and I published it: the AIO on-page checklist.
- Weekly measurement across all six models. One model is a sixth of the picture, and one check is a coin flip. The mention rate over weeks is the real number.
This is the exact loop the platform runs for the businesses in this post: fixed buyer prompts, six models, every week, with the misses turned into ranked actions. But the loop matters more than the tool. Whether you automate it or run it by hand on Friday afternoons, start measuring. The businesses in the 90% row aren't there because they're lucky. They're there because every question a buyer asks has a page, a fact, and a citation waiting to be quoted.
