How to check what ChatGPT says about your store, properly
Last reviewed
Short answer
Ask like a stranger, repeatedly. Use a temporary chat or an account with memory off, because OpenAI documents that Memory and custom instructions personalise answers, and your own account knows you. Run ten to twenty buyer-intent prompts, several times each across a week, and record mentions in a sheet. One run is noise. The pattern across runs is the measurement.
Why your own ChatGPT is the wrong instrument
OpenAI's Memory FAQ describes ChatGPT drawing on saved memories and chat-history insights to shape responses, and custom instructions apply across all chats. You have typed your store's name into that account for months. When it mentions your brand, you are not seeing what a shopper sees, you are seeing your own reflection.
The second trap is single runs. The best published measurement study, an April 2026 arXiv paper, found day-to-day similarity of cited sources averaging 0.34 to 0.42 and recommended at least seven runs per prompt for brand monitoring. The variance is the system working as designed, which deserves its own page. For the audit it means one answer, any answer, proves nothing in either direction.
The clean method, five steps
This is the manual version of what visibility tools automate, and it is free. Budget an hour for the first pass.
- Write 10 to 20 prompts a real buyer would ask: best X for Y, X under $80, is [competitor] worth it, where to buy X in [market]. Not your brand name, the category
- Open a temporary chat, or a clean account with memory and custom instructions off
- Run each prompt, record in a sheet: mentioned yes or no, position, what was said, who else was named
- Repeat the set two or three times across a week, fresh session each time
- Score it: mention rate per prompt across runs, plus the competitor names that keep recurring
Reading the results without fooling yourself
A 0% mention rate across three runs of twenty category prompts is a real finding: assistants do not currently reach for you, and the troubleshooter is the next page. A competitor appearing in most runs while you appear in none is also a real finding, and their reviews, mentions and product data are worth studying.
What is not a finding: appearing in one run and not the next, moving from sometimes to sometimes, or anything measured the week you changed nothing. The paper's advice is aggregation over weeks, and the honest unit is a trend line, not a screenshot.
When to stop doing this by hand
The method scales badly on purpose: three runs of twenty prompts is sixty answers to read, weekly, forever, and the value is in the forever. That cadence is what Entitled's visibility tracking automates, same sampling idea, weekly, scored for trend, framed as directional because that is what sampling honestly is.
Do the manual audit once first anyway. Reading sixty raw answers about your category teaches you things no dashboard summarises: the phrasing shoppers get back, which competitor claims recur, and what the assistants believe your products are for.
Frequently asked questions
Should I test on ChatGPT only?
Start there for reach, then run the same sheet on Perplexity and Google's AI Mode if your buyers use them. Keep the prompt set identical across engines so differences mean something.
What if ChatGPT says something false about my store?
Record the exact claim and check what the assistant could have read: your own pages, stale listings, old reviews. Wrong prices and dead products usually trace to stale data somewhere public. Fix the source, then re-test on the same cadence as everything else.
How many prompts is enough?
No vendor publishes a magic number. The published research recommends at least seven runs per prompt for stability, and commercial tools converge on tens to hundreds of prompts checked repeatedly. For a manual directional read, 10 to 20 prompts run a few times each is proportionate.
My brand shows when I ask with my name. Does that count?
Branded prompts test whether the assistant knows you exist, worth one row in the sheet. The commercial question is category prompts, where you compete for the recommendation. Score them separately and weight the category rows.
Sources
Keep reading
Score your own catalog
Entitled grades every product against the rules on this page, explains each failure, and fixes them with AI you approve. The score is free forever.
