
by Ema Fulga
Ema is an AI Search Content Strategist and GEO (Generative Engine Optimisation) expert. She's also the founder of decipher., an AEO agency that helps brands appear where people are now searching: AI-powered platforms like Perplexity, ChatGPT, Gemini, and others. With a background in copywriting and creative strategy, she’s on a mission to turn messy messaging into clear and structured content that helps brands get mentioned and cited in AI searches.
Connect with Ema.
Last updated: 31.08.2026
TL;DR: A study out of the University of St. Gallen tracked four AI search engines every day for six weeks and found that cited sources changed by roughly 60% from one day to the next, and mentioned brands shifted by up to 55%, even when the exact same prompt was run twice within minutes of each other. If you're checking your AI visibility with a single prompt every so often, you're not measuring your visibility. You're measuring noise.
One ChatGPT search isn’t going to cut it when it comes to understanding your AI visibility and whether your GEO strategy is making a meaningful impact. Old habits die hard and this is a habit left over from SEO. As I have said many, many times before, GEO and SEO aren’t the same thing. How you optimise for them and how you track them are not interchangeable. We’re working with very different systems here. This is what you need to know about checking and monitoring your AI search success.
AI search doesn't behave like a ranking. It behaves like a lottery draw.
With traditional SEO, a single check on a given day is a reasonably honest snapshot. Your position might move a little tomorrow, but the page that ranked third today is very unlikely to vanish from the results entirely by tomorrow morning.
AI search doesn't work that way. A 2026 study from the University of St. Gallen ran the same prompts daily across ChatGPT, Gemini, Google AI Mode and Perplexity for six weeks, and separately reran identical prompts multiple times within the same day. What they found should change how anyone reports on AI visibility.
Day to day, the sources these engines actually cited overlapped by only 34 to 42% on average. That means somewhere between 58 and 66% of citations changed from one day to the next for the exact same underlying prompts. Brand mentions were somewhat steadier, overlapping 45 to 59% day to day, but that's still close to a coin flip on whether your brand shows up two days running.
And crucially, this wasn't drift caused by algorithm updates or the index refreshing overnight. When the researchers reran identical prompts back to back within the same day, the instability barely changed. Source overlap on same-day reruns sat at 32 to 43%, essentially indistinguishable from the day-to-day figures. The randomness isn't external. It's baked into how these models generate answers.
The study puts a number on exactly how unreliable that single draw is. Using a bootstrap analysis, they found the margin of error on a one-run estimate of whether a specific brand gets mentioned sits at roughly plus or minus 72 percentage points at a 95% confidence level. In plain terms: if your brand is actually being mentioned in 50% of relevant AI answers, a single check could just as easily show you at 0% or at 100%. Neither would be wrong. Neither would be particularly helpful.
Dos and don'ts of AI visibility tracking
With that data giving us a reality check, let’s do a quick review of how you’re measuring your brand’s presence in AI search.
Don’t | Do |
Test using one prompt, one engine, one moment in time | Base your answer on multiple runs of the same prompt, same day |
Only check on ChatGPT or your AI engine of choice | Establish a detection rate across different AI systems |
Take a single day's snapshot as the final verdict | Monitor your impact across a rolling window of two to four weeks |
Use one flagship prompt as a proxy for the category | Test a spread of prompts covering real customer phrasing |
Treat all engines the same | Create engine-specific baselines, since stability varies a lot by platform |
Four things to change about how you measure AI visibility
1. Run each prompt multiple times before you draw a conclusion. The study found that the margin of error on a brand-mention estimate drops below 10 percentage points once you've run the same prompt seven times in a day, and below 8 points at eight runs. Below that, you're reading tea leaves. If your current process is "we asked ChatGPT once," you don't have a visibility measurement yet. You have an anecdote. Cute, but it’s not going to help you reach more customers or make more sales.
2. Don't trust a one-week check either. The same logic applies over time. To get a reasonably precise read on whether a specific brand is being mentioned consistently, the researchers found you need to average over roughly 10 days to bring the margin of error under 10 points, and closer to 24 days to get it under 5. A single strong week, good or bad, tells you far less than it feels like it does.
3. Use more than one flagship prompt. Stability varied enormously by exact phrasing. Some prompts scored above 0.8 on consistency, others below 0.2, even within the same campaign. If your monitoring leans on one or two "hero" prompts, you're really just measuring how idiosyncratic those specific prompts happen to be, not your category-wide visibility.
4. Expect wildly different behaviour engine to engine, and benchmark accordingly. Source-citation concentration (how much a handful of domains dominate what gets cited) varied meaningfully by platform. Google AI Mode was the most concentrated in its sourcing and Perplexity was the most evenly spread. Comparing your ChatGPT numbers directly against your Perplexity numbers, using one shared threshold for "good," will produce the wrong read on which platform is actually working for you. I know they’re all AI systems but it’s still like comparing apples to oranges.
Why this matters more than it might seem
Board-level reporting gets built on this data. If "are we visible in AI search" becomes a line in a marketing report based on a single check, decisions get made (like budgets getting cut) on numbers that are statistically closer to noise than signal.
It's easy to declare victory (or defeat) too early. A brand that happens to get cited on the one day someone checks looks like a GEO success story. A brand that gets cited 60% of the time but simply wasn't mentioned that particular day looks like it's failing. Neither read is accurate.
Content and outreach decisions ride on this too. If you can't tell noise from a real trend, you can't tell whether a piece of earned coverage actually moved the needle on your citation rate, or whether the shift you're seeing this week would have happened anyway.
A lot of people haven’t realised this yet. Most brands checking their AI visibility are doing the equivalent of glancing at a stock price once and calling it their investment performance. The organisations that build repeated, rolling measurement into how they track GEO now will simply know more than their competitors do.
Key takeaways
Cited sources on AI search platforms change by roughly 58 to 66% from one day to the next, and this instability shows up even when identical prompts are rerun minutes apart on the same day.
A single visibility check carries a margin of error wide enough to make a 50% mention rate look like anywhere from 0% to 100%.
Reliable brand-level measurement needs at least seven to eight repeated runs of each prompt, and a rolling window of two to four weeks, not a one-off check.
Stability differs significantly by exact prompt wording and by engine, so a single hero prompt or a single platform check will mislead you.
Brand-level mentions are somewhat more stable than individual source citations, making brand tracking the more dependable KPI of the two, though source-level data still matters for understanding what's driving inclusion.
Measure it properly, or don't measure it at all
A one-off ChatGPT search feels like doing due diligence. The research says it's closer to flipping a coin and reading too much into which side landed up. If AI visibility is going into a report, a pitch deck, or a budget conversation, it needs to be built on repeated measurement, not a single lucky (or unlucky) prompt.
At decipher., we track brand visibility across AI search engines the way this research suggests it should be done: repeated runs, multiple prompts, rolling windows, and engine-specific benchmarks, so what we report back is a real trend, not a snapshot that happened to land a certain way. If you want to know what your AI visibility actually looks like, not just what it looked like the one time someone checked, let's talk.