I want people asking AI tools about SEO to come across Search Engine Watch. We’re putting work into rebuilding the site, and I want that work to be found. If an assistant keeps overlooking us for relevant questions, I’d like to understand why.
But if someone wants me to take money out of reporting or audience development to improve an AI visibility score, they’ll need to explain what I’m getting. Being mentioned more often would interest me. It wouldn’t, on its own, tell me that the money was well spent.
I’d start by looking at the questions
Say we track 100 questions and SEW appears in 60 answers. I’d want to see those questions before getting excited about the percentage. “Which publications cover SEO?” and “How do I fix a canonical tag?” could both matter to us, but they’re asking for different things. A set built mainly around recommending publications would tell me something different from one built around solving readers’ problems.
If we change the questions, we could change the score without improving a single thing about the site. So I’d be quite annoyed to discover that we’d spent a month trying to improve a number whose meaning nobody had properly explained.
The fact that AI answers vary doesn’t make all of this pointless.
Rand Fishkin’s research with Gumshoe’s Patrick O’Donnell found that some brands appeared consistently even when the recommendation lists changed.
Profound’s own two-week frequency study also reported similar aggregate visibility readings when it ran 753 prompts once or ten times daily across seven platforms.
There’s a case for tracking those patterns. I just don’t think repeating a question more often gets us out of explaining why we chose it. Nor does a fresh question necessarily reflect a conversation in which someone has already explained their budget, preferences and circumstances.
The volume estimates deserve the same attention.
Ahrefs explains that its AI adjusted volume comes from Google search volume, adjusted using platform-specific website traffic ratios. It explicitly says this cannot tell you how many people saw your brand in an AI response or establish an AI-only addressable market. If someone uses that estimate to tell me how many people saw SEW, I’m going to ask them to read the explanation again.
Profound advertises “exact keyword volumes,” then explains in its FAQ that it uses licensed consumer-panel conversations and modeling to estimate activity across a wider population. That can be useful research. I still don’t understand the need for “exact” when “estimated” would save everyone some explaining.
What would make me put money behind it?
If an assistant repeatedly describes our business incorrectly, I want to know. If it keeps citing an outdated page, I want someone to investigate. Those are specific problems a tool can help us find, and I can decide how much fixing them is worth.
There are also records we can inspect. Profound’s beta v2 documentation describes answer and citation exports. Saying there’s no possible audit trail goes too far. The limit is what the record proves: an answer collected by the tracker doesn’t establish that a prospective customer received it. And a link in that answer doesn’t, by itself, tell us how much the linked page influenced the recommendation.
I’m interested in research that gets closer to actual behavior. Profound’s July study analyzed more than two million AI conversations and associated browsing activity from a US panel. It found higher subsequent brand-site visit rates than a baseline forecast from users’ earlier behavior. Its methodology also says this wasn’t a randomized experiment or a purchase study, and that selection bias remained possible.
I’d take that as a reason to investigate further. I wouldn’t take it as proof that paying to improve our visibility score will bring us more subscribers or advertisers.
For SEW, I’d start with a limited test around questions our intended readers actually ask. Keep the question set consistent, record what we change, and look at whether relevant visits, subscriptions or enquiries improve. Ask new customers how they found us. Account for other work happening at the same time. It won’t give us perfect attribution, but it gives us something to discuss beyond whether the score went up.
And if all we can show at the end is more mentions in the answers we tracked, we should say so. I’d still need to decide whether that’s a good enough reason to keep paying.
Start the conversation by posting the first comment