
TL;DR: A proper cross-engine visibility audit requires five specific steps run consistently: building a fixed list of 15 to 25 realistic buyer questions, running each one across ChatGPT, Perplexity, and Gemini separately rather than assuming consistent behavior across engines, recording not just whether you're mentioned but which competitors appear alongside you, repeating the same query set at least twice across separate sessions to rule out random variation, and logging every result in a structured format that supports comparison against the next audit cycle. Skipping any one of these five steps, especially running the same query only once or only on one engine, produces a result that looks like data but isn't reliable enough to act on.
Typing one question into ChatGPT, noting whether your company gets mentioned, and calling it a visibility check is the AI-era equivalent of checking website traffic by looking at the number once and never again. A proper audit requires structure: a fixed set of questions, multiple engines, repeated sessions, and a consistent way of recording what actually happened.
AI-generated responses vary between sessions, even for the identical query, and vary further between different engines, which use different underlying models and retrieval methods entirely. A single query run once on ChatGPT alone captures one data point out of a much larger space of possible outcomes, and treating that single point as representative of your overall visibility is a statistically weak basis for any real decision about content strategy or investment.
| Step | What it requires | Why it matters |
|---|---|---|
| 1. Build a fixed question set | 15 to 25 realistic buyer questions, written the way a real prospect would ask them | Ensures every audit cycle tests the same ground, making results comparable over time |
| 2. Run across all three engines separately | The same question set tested individually on ChatGPT, Perplexity, and Gemini | Each engine retrieves and cites differently; one engine's result doesn't predict another's |
| 3. Record competitor presence, not just your own | Noting every named competitor appearing in the same response | Your own presence or absence only means something in context of who else appears |
| 4. Repeat across separate sessions | Running the same query set at least twice, on different days | Distinguishes a stable pattern from a one-off session variation |
| 5. Log results in a structured, comparable format | A spreadsheet or similar record capturing engine, query, mention status, and competitors | Makes this audit comparable to the next one, turning a snapshot into a trend |
A company writing its own audit questions often unconsciously uses internal terminology or category language that doesn't match how a real prospect would actually phrase the same question. Writing the question set based on language pulled directly from actual sales call transcripts or customer discovery notes, rather than from internal positioning documents, produces a far more realistic test of what a real buyer's AI-assisted research would actually surface.
purple path's analysis of what ChatGPT and Perplexity actually pull from when answering a question covers exactly why these engines behave differently at a mechanical level, live retrieval versus training-data reliance depending on mode. Gemini adds a third distinct retrieval and ranking approach into the mix. Testing only one engine and assuming the result generalizes to the others isn't a reasonable shortcut; it's a different, considerably weaker audit that happens to look similar on the surface.
A "no mention" result for your own brand means something different depending on whether competitors are present in that same response. If no company at all is mentioned for a given query, that's a genuine whitespace opportunity. If two named competitors are mentioned and you aren't, that's active competitive displacement requiring a different kind of response. Recording only your own presence or absence, without also capturing who else appears, discards exactly the context needed to correctly interpret what any single result actually means.
AI model responses aren't perfectly deterministic; the same query run twice can produce somewhat different results even without any underlying change to the source content or the model itself. Running the full query set at least twice, ideally with a few days between sessions, and looking specifically for results that hold consistent across both runs versus results that flip between runs, distinguishes a genuinely stable pattern from noise that a single-session audit can't tell apart.
An audit run carefully but recorded only in scattered notes or memory produces no lasting value beyond the single session it was performed in. A structured log, one row per query per engine per session, capturing mention status and every competitor observed, is what allows this quarter's audit to be directly compared against next quarter's, turning a one-time snapshot into the kind of trend data covered in purple path's quarterly AEO audit template.
For a set of 20 questions across three engines, run twice across separate sessions, expect roughly 120 individual query-and-response reviews total. With a consistent, focused approach, this typically takes a half day to a full day of dedicated time, which is a meaningful but manageable investment for the quality of data it produces compared to a single, quick, unstructured check.
A person too close to a company's own content and positioning can sometimes read an ambiguous AI response more generously than an outside observer would, unconsciously crediting a vague or partial mention as a genuine positive result. Having someone less directly involved, or even someone entirely outside the company, review a sample of the raw AI responses and independently judge whether your brand was genuinely, clearly mentioned in a meaningful way adds a useful check against this kind of optimistic self-interpretation.
Beyond recording a simple mention or no-mention judgment for each query, saving the actual full text of each AI response provides a much richer record to revisit later, particularly useful when trying to understand exactly how a company was characterized, not just whether it was mentioned at all, or when a later dispute arises about whether a particular result should have counted as a genuine citation or a borderline, ambiguous mention.
Dedicated GEO measurement platforms can automate much of this process, running consistent query sets across multiple engines on a schedule and logging results automatically, which is worth considering once a company is committed to running this audit regularly rather than as a one-time exercise.
Fifteen to 25 questions covering the most realistic and commercially important buyer research topics tends to produce a reliable picture without becoming unmanageable; adding many more questions provides diminishing additional insight relative to the added time cost of running and reviewing them all.
The core set should stay largely consistent to support trend comparison over time, though it's reasonable to add a small number of new questions reflecting emerging topics or new competitors, while keeping the majority of the set unchanged for genuine comparability.
A result that flips between sessions suggests a less stable, more borderline citation status, worth flagging as an area to watch closely in the next audit cycle rather than treating either individual session's result as the definitive answer.
Yes, a smaller, more focused question set specific to a narrow category still provides valuable, actionable data; the five-step structure matters more than the absolute size of the question set, and a smaller set run rigorously beats a larger set run carelessly.
Running this five-step method for the first time builds the structured baseline every future comparison depends on. Talk to purple path about running a properly structured cross-engine visibility audit for your brand.

Dave leads purple path's content team, getting clients' inbound, outbound, thought leadership, social, and video content running fast, and making sure it actually works. In an AI-saturated content landscape, he's focused on the thing that still wins: content that engages and delivers real value.He's spent his career shaping content marketing strategy for SaaS companies globally, and previously as Head of Content at Minit Process Mining and Senior Copywriter at Exponea. He also built and exited his own company, Elite Language Center, over nearly nine years as CEO. His work has been featured in Forbes, and he's increasingly focused on LLM visibility, making sure content shows up where AI-driven search is heading next (GEO/AEO).