To check your brand's AI visibility yourself, write a fixed list of 20 to 30 prompts a buyer would type, run each one three times in ChatGPT, Perplexity and Google with personalization switched off, and log every answer in one sheet. Repeat it weekly for 30 days. The result is a dated baseline, not a ranking.
Most brands skip this step and go straight to fixes, which leaves no way to tell later whether the fixes changed anything. A baseline built by hand costs a few hours a week and no software. It will not tell you how often real customers see your brand, because no outside test can measure that. It tells you something narrower and still useful: across a defined set of questions, on known dates, how often each engine mentions you, whether what it says about you is correct, and which sources it leans on. This guide gives the protocol, the log template and the rules for reading the numbers at the end of the month.
Key Takeaways
- Split prompts into a discovery set that never names you and a brand set that does, and freeze the wording for 30 days
- Run each prompt three times per engine, because a single run cannot separate a real absence from a bad draw
- In ChatGPT, use a temporary chat set to Unpersonalized before the first message
- Log one row per run and keep the unfavorable runs, or the baseline stops meaning anything
- Search Console's generative AI report adds impressions; utm_source=chatgpt.com adds ChatGPT referrals
- Report mention rate with its denominator, and treat it as a sample, never as a position
Which prompts should go into an AI visibility check?
Split the prompts into two sets and keep them apart for the whole 30 days. The brand set names you: what does the company do, what does it cost, who are its competitors, how do I hire it. The discovery set never names you and reads like a buyer who has not heard of you yet: best agency for a given problem in your city, how do I fix a specific issue, what should I look for in a provider. Discovery prompts matter commercially, because they show whether an engine brings you up without being asked. Brand prompts matter for accuracy, because they show whether an engine describes you correctly when someone already knows your name. Mixing the two sets into one score hides both answers.
Write the prompts the way customers actually type them. Pull the wording from sales call notes, support tickets, the questions prospects ask on a first call and the queries in your own Search Console. Google's autocomplete is a free way to confirm that a phrasing exists, although it gives no volume. Fifteen discovery prompts and eight to ten brand prompts are enough for a first baseline. Never ask an engine to recommend your brand or to compare it favorably, since a leading prompt measures the prompt and not the market. Once the list is written, freeze it: the same text, word for word, every week. Changing a prompt in the middle of the month breaks the comparison for that row.
How do I stop personalization from skewing the results?
Personal history changes answers, so the test has to remove as much of it as each product allows and record what it could not remove. In ChatGPT, OpenAI's help center says a temporary chat can still use existing memories, custom instructions and plugins unless you choose Unpersonalized before the first message, and that the choice cannot be changed once the conversation starts. Unpersonalized does not use memory, custom instructions or plugins, and a temporary chat does not create or update memories. Temporary plus Unpersonalized is therefore the cleanest setting for the log. Write down the model name shown on screen and whether the answer used search, because both change what comes back. Source: OpenAI Help Center, Temporary chat in ChatGPT, accessed October 2026.
For Google, run the prompts in a private browser window, signed out, from the country you sell in, and record whether an AI Overview appeared at all, since many queries show none. Do the same in AI Mode. In Perplexity, start a fresh thread for every run. Perplexity's help center says each answer includes numbered citations linking to the original sources, which makes it the easiest engine to log sources from. Whatever you cannot control, such as location or a model update in the middle of the month, goes into the notes column with the date. An honest note about a confound is worth more than a clean number that hides it. Source: Perplexity Help Center, How does Perplexity work, accessed October 2026.
What should an AI visibility prompt log record?
The log takes one row per run, and each prompt produces several runs. If a prompt runs three times in three engines, that is nine rows, and every one of them stays in the sheet, including the runs where you do not appear. Deleting unfavorable runs is the fastest way to build a baseline that lies. The table below is the template, column by column. Copy it into a Google Sheet, freeze the header row and add a dropdown to the yes or no columns, so that three people logging on different days still write the same values. Keep the full answer text in its own column or as a linked screenshot, because a summary written on the day loses the details you will want later.
| Column | What to write | Why it matters |
|---|---|---|
| Run date and time | Date, time and time zone | Answers change; the date is the evidence |
| Engine and mode | ChatGPT, Perplexity, Google AI Overview, Google AI Mode | Each engine is a separate baseline |
| Model shown | The model name visible on screen, or not shown | A model update can move results overnight |
| Prompt ID and set | D01 to D15 for discovery, B01 to B10 for brand | Keeps the wording frozen and comparable |
| Repetition | 1, 2 or 3 | Shows how much answers vary between runs |
| Search used | Yes, no or unclear | An answer without search reflects the model, not your current site |
| Brand mentioned | Yes or no | The numerator for mention rate |
| Described correctly | Yes, partly, no, or not mentioned | Appearing with wrong facts is its own problem |
| Competitors named | Names, in order of appearance | Shows who owns the answer today |
| Sources cited | URLs or domains | Shows which pages the engine trusts for the topic |
| Your URL cited | Yes or no, and which page | Separates a mention from a citation |
| Full answer | Text or a screenshot link | The raw evidence behind every other column |
| Notes | Anything you could not control | Location, login state, outages, model changes |
How often should I run the prompts over 30 days?
Five runs across the month give a usable baseline without taking over the calendar: day 1, day 8, day 15, day 22 and day 29. Each run covers every prompt, three repetitions each, in every engine you track. With 25 prompts and three engines that is 225 rows per run, which sounds heavy, and most of the time goes into copying answers rather than reading them. If the load is too much, cut engines before you cut repetitions, since one repetition per prompt cannot tell a real absence from a bad draw. Keep the same day of the week and a similar time of day for every run, and log anything that changed on your own site during the month, such as new pages, a redesign or a robots.txt edit.
What do Search Console and analytics add to a manual baseline?
The prompt log tests a sample, while Search Console and analytics count real activity, and the two answer different questions. Google's help page for the generative AI performance report says it shows impressions, meaning how many times links to your site were shown in a generative AI feature on Google Search, and that it covers AI Overviews and AI Mode. Its dimensions are pages, countries, dates and devices, with a filter for text-based and multimodal search, and it excludes Search Labs experiments. Google says the insights reached all websites worldwide on August 31, 2026, and that a site may not see the report if it has too few generative AI impressions. Export the 30 days that match your log. Source: Google Search Console Help, accessed October 2026.
On the ChatGPT side, OpenAI's publishers FAQ says ChatGPT automatically adds utm_source=chatgpt.com to referral URLs, so visits from ChatGPT search can be counted in Google Analytics or any tool that reads UTM parameters. Build a saved report filtered on that source and, where your analytics shows them, on referrals from perplexity.ai, then export the same 30 days. Expect small numbers and some leakage, because not every visit influenced by an AI answer arrives with a referrer, and some people read an answer and search for the brand by name days later. The analytics column sits next to the prompt log and never replaces it: the log shows what the engines say, and analytics shows part of what people did next. Source: OpenAI Help Center, Publishers and Developers FAQ, accessed October 2026.
How do I read the baseline after 30 days?
Report mention rate per engine and per prompt set, with the denominator next to it: valid runs where the brand was mentioned, divided by all valid runs, written as a fraction such as 7 of 45 so the reader sees the sample size. Report discovery and brand prompts separately, and do the same for accuracy: of the runs where you appeared, how many described you correctly. Then read the competitor and source columns, which are usually more useful than your own score, because they show which domains and which kinds of pages the engines rely on for your category. A competitor cited from a directory, a comparison article or a forum thread points to a different fix than one cited from its own pricing page.
Three limits belong at the top of the report, where the reader sees them first. A three-repetition sample is exploratory and is not an estimate of what the market sees. A mention rate is not a ranking, and no tool can give you a fixed position in ChatGPT, because answers are generated per conversation. A change after the first month is a correlation until you have ruled out what else moved, such as model updates, seasonality, your own site changes or a competitor's launch. The baseline exists to make the next decision honestly: which missing pages to write, which wrong facts to correct at the source and which crawler and eligibility checks to run first, in the order we lay out in our GEO audit checklist.
Common questions
How many prompts do I need for an AI visibility baseline?
Twenty to thirty is enough for a first baseline: about fifteen discovery prompts that do not name the brand and eight to ten brand prompts that do. Each one runs three times per engine, because answers vary between runs. More prompts widen coverage, but the wording has to stay frozen for the whole period or the comparison breaks.
Should I be logged in to ChatGPT when I test?
You can be logged in, as long as you open a temporary chat and choose Unpersonalized before the first message. OpenAI's help center says Unpersonalized does not use memory, custom instructions or plugins, and that a temporary chat does not create or update memories. A regular chat follows your account's personalization settings, which can skew the result.
Does the Search Console generative AI report show which prompts triggered my impressions?
Google's documentation lists four dimensions for the report, pages, countries, dates and devices, and it does not list queries or clicks. The report tells you that your pages appeared in AI Overviews or AI Mode and where, which complements a manual prompt log and does not replace it.
How do I track ChatGPT traffic in Google Analytics?
OpenAI's publishers FAQ says ChatGPT automatically adds utm_source=chatgpt.com to referral URLs from ChatGPT search. Filter sessions on that source in Google Analytics and export the same date range as your prompt log. Some visits influenced by an AI answer arrive without a referrer, so read the number as a floor.
Can a tracking tool replace the manual prompt log?
A tracking tool can run more prompts more often, and it is worth paying for once you know which prompts matter. It still needs a frozen prompt set, a record of the model and settings used, and someone reading the answers for accuracy. Running the log by hand for the first month tells you what to ask a tool for.
Sources
- Temporary chat in ChatGPT, OpenAI Help Center, accessed October 2026
- Publishers and Developers FAQ, OpenAI Help Center, accessed October 2026
- Generative AI performance report (Search), Google Search Console Help, accessed October 2026
- How does Perplexity work?, Perplexity Help Center, accessed October 2026
Want the baseline run for you?
We run the same protocol with a frozen prompt set and hand over the full log, including the runs where you do not appear. We do not sell guaranteed citations.
Request a GEO audit