At the start of summer (~3 months ago), I built my own LLM visibility tracker using Claude Code.
The idea was simple: Track the AI-generated answers and citations across the four major LLMs (ChatGPT, Claude, Gemini & Perplexity) for 12-40 prompts per client, run weekly.
I also tracked the Google SERP results with similar or identical keywords to compare traditional Google results to AI chatbot answers. I wanted to understand what each engine was citing, how the answers and citations varied across LLMs, and how they changed over time to answer a few simple questions:
- Do the different LLMs agree on who to cite? Do they agree on who to recommend?
- How do citations and recommendations change over time, from one run to the next?
- Do the Google SERPs line up with AI-generated answers?
The purpose was three-fold; to play around with building my own tools with AI, to track AI visibility for my SEO clients, and to gather some data on what’s working right now in AI search.
I’ve recorded 13,184 citations and 1,765 answers so far. This article aims to break down what I’ve learned.
Methodology: How I collected this data
The setup is simple enough to replicate. I asked Claude Code to help me build an LLM citation tracker that scrapes the text answer and logs every URL each model cites, plus whether each brand gets mentioned in the answer text, and runs every week. Then I did some testing and tweaking till I liked the result.
I tracked 72 prompts across three B2B brands:
- My own website (billwidmer.com)
- My friend’s Amazon agency (evolveadagency.com)
- My client, Semrush (semrush.com)
I included questions like “best all-in-one SEO platform for marketing teams” and “is it worth paying an agency to manage Amazon?” Every week, a script runs each question through four models: ChatGPT (with web search forced on), Claude, Perplexity, and Gemini.
The citations and answers went into a spreadsheet:

Now, there are some limitations to this data I’m aware of:
- 12 to 40 questions per brand is a directional sample, not statistical proof.
- My Claude data comes from an agent-with-search setup rather than the raw API, so its citation style may differ slightly from claude.ai. The other three are run with API.
- Gemini sometimes only reveals the domain it cited, not the exact page.
Data through August 23, 2026.
A note on mentions vs. citations
For clarity’s sake: A “mention” is when the AI names or recommends your brand in its answer. A “citation” is when it links your page as a source.
In my data, citations are almost a strict subset of mentions: out of 1,052 query runs, only 17 cited a brand’s page without also mentioning the brand. Mentions without citations, though, are everywhere.
Across the whole dataset:
- Claude: brands mentioned in 70% of answers, cited in 40% (a 30-point gap)
- Gemini: 72% mentioned, 41% cited (31 points)
- Perplexity: 67% mentioned, 44% cited (23 points)
- ChatGPT: 56% mentioned, 47% cited (only 8 points)
ChatGPT is the outlier. When it names a brand, it usually links the brand, too. Claude and Gemini talk about brands freely without sourcing them.
Finding #1: The models almost never cite the same sources
Across 1,792 query-and-domain combinations, all four models only agreed on citing the same domain for the same question 30 times. That’s 1.7%.
Even the two biggest names barely overlap. On one brand’s question set, ChatGPT and Perplexity shared 7.6% of their cited domains. On another, 5.4%.
Each model appears to have a consistent sourcing personality:
- Perplexity cites heavily: 19.2 sources per answer on average. Its all-time favorites in my data are LinkedIn (374 citations), YouTube (248), and Reddit (232).
- Gemini averages 8.1 sources per answer and leans on Reddit and listicle roundups.
- ChatGPT averages 4.5 and skews towards primary sources like Google’s own documentation, arXiv papers, and vendors’ pricing and help pages.
- Claude averages 3.6 and loves review sites and curated listicles. The review platform Clutch is its single most-cited third-party domain in my data.
The most interesting finding, in my opinion, is that across hundreds of answers, Claude never cited Reddit once.

My data backs him up. In last week’s run across 72 B2B buyer questions, Reddit was 0% of Claude’s citations, 0% of ChatGPT’s, 3% of Perplexity’s, and 1% of Gemini’s.
This also tracks with ChatGPT’s recent update that slashed Reddit citations to near-zero.
What got cited instead? Niche listicles, review platforms, and comparison posts.
Finding #2: The models don’t agree with Google, either
Ranking on page one of Google is neither necessary nor sufficient for getting cited by AI. For the questions in my tracker where I could map the conversational prompt to a keyword search, I pulled Google’s top 10 organic results the same week and measured overlap with what each model cited.
- At the domain level, only 10% to 30% of the domains the models cited also ranked in Google’s top 10 for the matching query, depending on the model and question set.
- At the exact-URL level, it collapses: roughly 1% to 5% for Gemini and ChatGPT. Perplexity tracks Google closest, and even it only overlaps 13% to 23% at the URL level.
Growth consultant Lara Stiris told me a story that backs this up: Researching accounting software for her own business, she asked AI for recommendations and the biggest incumbent in the category never came up.
“All that came up was startups… companies that were newer, that had more social interactions around them,” she said. A brand with enormous Domain Authority and a mountain of ranking content, absent from the conversation. “That to me is entirely different from SEO.”
Before you conclude SEO is dead (it isn’t), here’s the balancing data point: Fruzsi Peti, a startup marketing consultant I interviewed, told me her clients earned AI visibility without any dedicated AEO budget.
“We have been focusing on PR activities… and what we noticed is that that kind of activity already put us forward in agentic search.”
Strong fundamentals still feed the machine. The assets overlap even when the URLs don’t.
Finding #3: The models don’t even agree with themselves
I ran the same 72 questions through the same models twice, two days apart. Claude kept only 30% of the URLs it had cited two days earlier. Gemini kept 38% of its domains. Perplexity was the most stable at 65%.
In Andy’s off-site optimization piece, Britney Muller cautioned that AI outputs are non-deterministic and that even ten runs of a prompt give you “a very crude directional signal.”
The takeaway? A single AI answer is weather. A tracked rate over weeks is climate. Never make a long-term strategic decision (or panic) based on the weather.
Bonus: The more famous you are, the less your citations tell you
One interesting thing I noted is that Semrush, which has a huge online footprint, had a 50-point gap between mentions and citations. Compare that to my website and Evolve Ad Agency, which had a 5- and 4-point gap (as of last week’s run).
Semrush (an established SEO software company) gets mentioned in 83% to 100% of relevant answers on every model. Its citation rate on those same answers ranges from 83% all the way down to 0%. In last week’s run, half of its query runs were mention-without-citation.
Here’s my theory: Famous brands live in the models’ training data, so the AI recommends them from memory while sourcing the answer from third-party pages. Unknown brands only get mentioned when the model actually retrieves and reads their site.
Granted, Semrush has far more content than either of the other two sites, so this could also just be that my brand and Evolve’s brand were only mentioned when their name was in the prompt. But this could also back up the theory that AI doesn’t pull citations when it has sufficient training data.
Which leads to a rule I haven’t seen anywhere else: citation dashboards understate visibility for big brands and track it almost one-to-one for small ones.
A 0% citation week means something completely different for a household name than for a startup. Know which end of that spectrum you’re on before you panic, and before you let a vendor scare you with it.
Conclusion: You don’t need four different strategies. You need one comprehensive one.
After all this divergence data, you’d expect me to prescribe a per-engine playbook. Reddit for Gemini, LinkedIn for Perplexity, etc.
Nope! I actually think the answer is creating a single, comprehensive marketing strategy. My belief is that AEO/GEO is not just SEO. It’s plain ol’ good marketing.
But don’t take my word for it. I interviewed Nick Lafferty, Founding Marketing Engineer at Profound, the AI visibility platform. When I asked whether they optimize differently for different engines, given the variation in their own data, he was blunt:
“Candidly, we’re not doing things differently for different engines. We’re kind of just optimizing for AI search broadly as a bucket.” Their priority is doing genuinely interesting things (original research, a podcast, even subway ads), knowing the AI visibility follows: “It’s like, let’s do this thing because it’s cool. We think our audience will like it. And then it’ll have knock-on benefits in AI search.”
Lara Stiris pushes the same direction from the measurement side. She warns that people treat AI visibility data “as the absolute truth versus directional, because they don’t always understand the methodology underneath it,” and that plenty of vendors are “overselling what they actually do.”
So here’s how I reconcile my divergence data with their engine-agnostic advice: the work itself (real expertise, consistent entity information, presence on the sources that keep showing up) is one strategy that feeds all four engines at different rates.
The 4-step playbook
Here are the four things I’m focused on based on my findings:
- Build a list of real buyer questions and track it weekly. Pull questions from sales calls and support conversations, not keyword tools. Andy’s sales-call transcript method is the best version of this I’ve seen. Track both mentions and citations, per model.
- Find your category’s gatekeepers in the citation data. List every third-party page cited instead of you, count repeats, and pitch the ones that keep appearing. One inclusion on a page that all four models trust beats four engine-specific tactics.
- Judge trends on three-plus weeks of data. With 30% to 65% week-to-week churn, single snapshots will make you chase ghosts.
- Keep doing the fundamentals. The engines disagree on URLs, but they all reward the same underlying things: real expertise, consistent entity information across the web, and third-party validation. That’s why one strategy can serve four disagreeing machines.
At the end of the day, this is a fun experiment I ran. If you notice any other flaws or gotcha’s, DM me on LinkedIn and I’ll try to address them!


