Someone asked ChatGPT for a business like yours this week. You have no idea what it said.
There’s no impression count for that, no ranking drop, no notification. The prospect got two or three names, called one, and you never knew you were in the running. The only way to find out is to go and look — properly, not once.
Most businesses do look, eventually. They open ChatGPT, type their service and their suburb, see a competitor, and either panic or shrug. That single check is worth almost nothing, and this article explains why, then gives you a test that is.
What is AI visibility tracking?
AI visibility tracking measures how often AI assistants name or link your business when someone asks a question you should own. Unlike keyword rank tracking, it produces a rate rather than a position, because the same prompt returns a different answer each time. A useful measurement needs a fixed prompt set, repeat runs, and a record of every answer.
That last part is where most attempts fall over. A screenshot is an anecdote. A rate is a metric. The difference matters because you can only act on something that moves in a direction you can defend.
It sits alongside, not instead of, your existing reporting. If you already watch a search visibility score for organic rankings, treat this as the parallel number for generative surfaces — measured differently, because the surfaces behave differently.
Why one ChatGPT check tells you nothing
AI answers are generated, not retrieved. The model samples from probabilities at each step, so two identical prompts produce two different answers — different businesses, different order, different list length.
The scale of that instability is now measured rather than assumed. SparkToro and Gumshoe.ai had 600 volunteers put 12 identical prompts through ChatGPT, Claude and Google’s AI around 2,961 times, normalising every response into an ordered list of brands. Ask for recommendations in a category a hundred times and your odds of seeing the same list twice are worse than one in a hundred — the list, the order and even the number of items all shift, sometimes returning two or three names and just as often ten or more. MediaPost Publications + 2
Day-to-day churn is more contained but still real. One 90-day study re-ran 1,247 buyer-intent prompts daily from March to May 2026 across eight platforms, diffing 897,840 answers, and found roughly 17% of prompts return a different set of recommended brands than the previous day, with the median brand list surviving unchanged for five days — three on Perplexity. MaxAEO
The same study found something more useful for planning. On ChatGPT, an average of 2.3 brands per prompt appeared in at least 80% of daily snapshots — the anchor brands — while the remaining slots rotated among four to nine challengers. MaxAEO
That’s the whole game in one sentence. There are roughly two anchor slots per question and everything else is a lottery. Your job isn’t to appear once. It’s to become an anchor on the few questions that pay.
Run every test signed out, in a private window, with chat history and memory turned off. Testing from your own logged-in account contaminates the result — the model has already watched you look up your own business, and the answer you get back is not the answer a stranger gets.
How many times do you need to run the same prompt?
Three runs per prompt tells you whether your business is absent, occasional, or reliably named — nothing more precise than that. Ten runs is the minimum for a percentage, and even then a three-out-of-ten result means the true rate sits somewhere between 11% and 60%. Reserve larger samples for the handful of prompts tied to revenue.
Here’s the arithmetic, because it’s rarely shown. These are 95% confidence intervals for an observed 30% mention rate at different sample sizes:
| Runs of the prompt | Times you were named | Observed rate | True rate is somewhere between |
|---|---|---|---|
| 10 | 3 | 30% | 11% and 60% |
| 30 | 9 | 30% | 17% and 48% |
| 100 | 30 | 30% | 22% and 40% |
| 10 | 0 | 0% | 0% and 28% |
Read the last row twice. Zero mentions across ten runs does not prove you’re invisible — your true appearance rate could still be as high as one in four. It does prove you’re not an anchor, which is enough to act on.
It’s also why an agency reporting “your AI share of voice is 23%” without naming the sample size and the prompt set is reporting a feeling. Ask for the denominator.
The 10-3-3 baseline: a test you can run this week
Ten prompts. Three surfaces. Three runs each. Ninety answers, two to three hours of work, and a number you can compare against next month.
Step 1 — Write ten prompts the way a customer would
Not keywords. Prompts. Nobody types “plumber Wollongong” into ChatGPT; they type “my hot water system is leaking, who should I call in Wollongong on a Sunday”.
Build your ten from four types:
- Direct recommendation. “Who are the best [service] providers in [suburb]?” The money prompt, and the hardest to win.
- Problem-first. “My [problem]. What are my options and who fixes this near [suburb]?” Customers describe symptoms, not services.
- Comparison. “[Your business] vs [competitor] — which is better for [use case]?” Tests whether the model can describe you at all.
- Brand check. “What do you know about [your business name]?” If it gets your services, location or ownership wrong, that’s a data problem rather than a content problem, and it’s usually the fastest thing to fix.
Include at least two prompts carrying a suburb or region. Australian search behaviour is heavily local, and generic national prompts will flatter or bury you for reasons that have nothing to do with your actual market.
Step 2 — Pick three surfaces that matter in Australia
Test where Australians actually are, not where the American blog posts point. Adoption here is high and lopsided: Roy Morgan’s March-quarter 2026 research put 13.6 million Australians aged 14 and over — 58% of that population — using AI tools in an average four-week period, with ChatGPT leading at 10.5 million. Telsyte’s 2026 study counts higher across the board: ChatGPT 13.8 million, Gemini 9.1 million, Meta AI 5.6 million and Copilot 5.4 million. The two firms disagree on absolute numbers and agree on the order, which is what matters here. Windows ForumOptimising
| Surface | Australian reach (2026) | How it builds an answer | How to test it |
|---|---|---|---|
| ChatGPT | 13.8M users (Telsyte); 67.7% of Australian AI chatbot share (Statcounter, June 2026) Optimising | Training data plus live web retrieval | Temporary chat, signed out, memory off |
| Google AI Mode & AI Overviews | AI Mode switched on in Australia on 8 October 2025; Olivetree Marketing AI Overviews since October 2024 | A customised version of Gemini, with follow-ups and multimodal input, SmartCompany fanning queries across Google’s index | Incognito, google.com.au, both the AI Mode tab and standard results |
| Microsoft Copilot | 5.4M users; 14.2% of Australian chatbot share, now clearly second Optimising | Bing’s index | copilot.microsoft.com, signed out |
| Gemini (app) | 9.1M users but only 7% of chatbot traffic share Optimising | Google’s index | Skip the app — most Gemini exposure reaches Australians inside Search |
| Perplexity | 5.1% of Australian chatbot share Optimising | Heaviest live retrieval of the group | Worth one pass; it reflects site changes fastest |
For most Australian businesses the three to test are ChatGPT, Google AI Mode and Copilot. The Gemini row is what trips people up: high user count, low direct traffic, because Australians meet Gemini inside Google Search rather than at the Gemini app. Test Google’s surfaces, not Google’s chatbot.
Step 3 — Run each prompt three times and score in bands
Three runs won’t give you a percentage. It gives you a band, and a band is a decision.
| Band | Result across 3 runs | What it means | What to do next |
|---|---|---|---|
| Absent | 0 of 3 | The model either doesn’t know you or can’t describe you confidently enough to name you | Fix entity fundamentals first — Google Business Profile, NAP consistency, LocalBusiness schema |
| Contested | 1–2 of 3 | You’re in the rotating tail, competing for a challenger slot | Build corroboration: third-party mentions, reviews, citations on sites that aren’t yours |
| Anchor | 3 of 3 | You’re one of the two-ish stable names for this question | Defend it. Keep the page fresh and re-test monthly for drift |
Score mentions and citations separately. A mention is the model naming you in the answer text. A citation is it linking your domain as a source. You can be mentioned without being linked — perception without clicks — or cited without being recommended. Tracking one hides half the picture.
Step 4 — Deepen on the prompts that pay
Pick the three prompts with the clearest commercial intent, usually the direct-recommendation ones, and run those ten times each on your single most important surface. That’s 30 more answers, and it gives you a rate you can trend with the confidence interval from the table above attached to it.
Everything else stays at three runs as a screening pass. Screening wide and deepening narrow is the only version of this a small team sustains past month two.
What counts as good AI visibility for a local business?
Good AI visibility for a local business means being named in most answers for the three or four prompts that actually drive enquiries, not appearing everywhere. AI answers typically name one to three businesses, so the realistic target is anchor status on a narrow prompt set rather than a high average across fifty prompts.
Chasing a broad average is how agencies produce impressive reports and no enquiries. Ten prompts at 40% each looks better on a dashboard than two prompts at 90%, and is worth considerably less.
Never report a percentage from a small sample without the sample size beside it. “Named in 8 of 30 runs across 3 prompts, ChatGPT, week of 10 August” is a sentence you can still defend in twelve months. “27% AI share of voice” is not.
Wiring it into reporting you already run
The manual test tells you what AI says. Three existing sources tell you whether it’s translating into anything.
- Segment AI referrals in GA4. Build a comparison on session source containing
chatgpt.com,perplexity.aiandcopilot.microsoft.com, save it, check it weekly. Volume will be small. The trend is the point, and these visitors usually arrive further along than organic ones. - Watch branded search in Google Search Console. People see you named in an AI answer, then Google your business name to check you’re real. Rising branded impressions against flat non-branded impressions is one of the few reliable signals that AI mentions are landing.
- Annotate your dates. Record when you ran each baseline and when you shipped each fix. Without dates you can’t separate a real improvement from the five-day churn the volatility data predicts.
None of these replaces the manual test, because a large share of AI answers produce no click at all. Being named and not clicked still moves the enquiry — it just doesn’t move your analytics.
Where the DIY method breaks down
Be honest about the limits before you build a quarter’s strategy on ninety answers.
- You’re testing one location. Answers vary with where the request appears to come from. Testing from your office tells you about your suburb, not your whole service area.
- Three runs is a screen, not a measurement. Treat every band as provisional until you’ve deepened.
- Prompt phrasing moves the result more than you’d like. Real customers phrase things a hundred ways you didn’t write down, so your ten prompts sample intent rather than census it.
- It doesn’t diagnose. The test tells you that you’re Absent. It doesn’t tell you whether that’s a schema problem, a review problem, or a “no third-party site on the Australian web has ever mentioned this business” problem. That diagnosis is what a proper GEO audit is for.
Below about five prompts and one surface, don’t formalise any of this. Check occasionally and put the time into your local SEO fundamentals instead, which is what feeds the models anyway.
What to do with what you find
The work that fixes an Absent result is mostly work you already recognise. Complete and accurate Google Business Profile. Identical business name, address and phone number everywhere. LocalBusiness schema that matches your visible page content. Steady, real reviews. Pages that answer a specific question in one self-contained place rather than burying the answer in paragraph nine.
What’s changed is the target. Generative engine optimisation aims at being the quoted source inside an answer rather than the tenth blue link under it, which rewards content that can be lifted off your page and still make sense with your name attached to it.
Block out three hours next week. Write your ten prompts, run them signed-out across ChatGPT, Google AI Mode and Copilot, and score each one Absent, Contested or Anchor in a spreadsheet with the date at the top. By the end of the afternoon you’ll know which of your service lines AI has never heard of — and that list, not a dashboard percentage, is your work for the quarter.







