
The 60-second answer
Probably not on day one. First establish a repeatable baseline around the questions that precede a purchase. Record whether your company is eligible to be found, mentioned, recommended, cited and - when somebody clicks - responsible for a qualified visit.
Search Console, Bing Webmaster Tools, analytics and controlled manual tests can reveal parts of that picture without a specialist subscription. A paid tool becomes useful when repeated collection, multiple markets, competitor tracking and reporting cost more to run manually than the subscription saves.
It should also expose the questions and answers behind its score. Otherwise, you have bought a dashboard before defining the decision.
An AEO tool can help you observe a market. It cannot guarantee that an assistant will recommend you.
In this article
- The five layers hidden inside AI visibility
- How to establish a useful baseline without an AEO tool
- What paid AI-visibility tools actually add
- When to stay manual, subscribe or commission an audit
- Seven questions to ask before buying
- What to fix after the measurement
1. The five layers hidden inside “AI visibility”
Here, an AEO tool means software that monitors how a brand or its pages appear in AI-generated answers. “Are we visible?” sounds like one question. It is at least five.
| Layer | What you are checking | What it does not prove |
|---|---|---|
| Eligibility | Can the relevant system access and use the page or business information? | That it retrieved, cited or recommended the company |
| Mention | Does the company appear anywhere in the answer? | That it is presented positively or as a suitable option |
| Recommendation | Is the company proposed for the buyer’s stated need? | That the supporting facts are correct or that your site is cited |
| Citation | Is your page used as a visible source? | That the company itself is recommended or that every linked claim is supported |
| Qualified referral | Did a person arrive and take a useful next step? | That the AI interaction caused the entire buying decision |
These layers can move independently. An assistant may cite your research while recommending a competitor. It may name your business without linking to it. A visitor may arrive from ChatGPT after several earlier interactions that analytics cannot see.
Tool vendors also use different definitions. Ahrefs separates mentions, citations, estimated impressions and AI Share of Voice. Semrush presents its own visibility score alongside mentions, citations and cited pages. Both can be useful within their declared methods; the headline numbers are not interchangeable.
Start with the layer your decision requires. Inspect answers to correct product facts, citations to understand source use and attributable visits to assess website actions - while keeping the attribution limit visible.
2. How to establish a useful baseline without an AEO tool
A manual baseline is not a free version of an enterprise platform. It is a smaller experiment designed around a real decision.
Choose questions from the buying process
Collect questions from sales calls, proposals, support conversations and Search Console. Include situations in which a buyer:
- defines the category;
- compares approaches or suppliers;
- introduces a constraint such as market, budget, integration or regulation;
- checks risk or credibility;
- asks for a shortlist.
Avoid variations invented only to make the company appear. OpenAI’s evaluation guidance recommends task-specific tests that reflect real-world use, logged evidence and human judgement. It addresses AI applications, but the discipline transfers: define the decision, then choose tests that represent it.
Record the conditions
For every test, retain:
- the exact question;
- the assistant or search product;
- the market and language;
- the date;
- whether the session was new or already personalised;
- the full answer and visible sources;
- the rule used to label mention, recommendation and citation.
Generated answers can vary. Repeat decision-critical questions and report the observed range. There is no universal responsible number of prompts or repetitions; the method must justify its coverage.
Use first-party signals where they exist
Google says the same SEO foundations apply to AI Overviews and AI Mode. Eligible pages must be indexed and able to appear with a snippet, but there is no special AI schema or extra technical requirement - and eligibility does not guarantee inclusion. AI-feature traffic is included in Search Console’s Web performance data rather than a complete cross-platform AI report (Google Search Central).
Bing’s AI Performance preview can show citations, cited pages, sampled grounding queries, URL-level activity and trends across supported Microsoft AI experiences. Bing says these measures do not indicate answer placement, page importance, authority or ranking (Bing Webmaster Tools).
OpenAI says publishers can allow OAI-SearchBot to access content intended for ChatGPT search. ChatGPT referral links include utm_source=chatgpt.com, allowing attributable visits to be analysed (OpenAI publisher guidance). That records clicks, not unclicked mentions or private conversations.
Together, these sources provide pieces of the baseline. None is a universal scoreboard.
3. What paid AI-visibility tools actually add
The strongest reason to buy a tool is operational reliability, not access to a secret optimisation method.
A suitable platform may add:
- scheduled collection instead of manual checks;
- consistent storage of questions, answers and sources;
- coverage across more assistants, countries or languages;
- competitor comparison under the same protocol;
- history that makes movement easier to investigate;
- exports and access controls for several stakeholders;
- alerts when an important answer or source changes.
These capabilities do not remove the need to choose representative questions, define counted events or inspect answers. Two tools may use different question databases, engines, locations, dates, entity definitions and formulas. Their scores are compressed views of different methods - not fixed ranks across the AI market.
Automation also scales mistakes. A biased question set remains biased; an ambiguous brand can distort comparison; a tool that treats every mention as success can reward an inaccurate recommendation.
4. When to stay manual, subscribe or commission an audit
Use the lightest method that can support the decision.
| Situation | Sensible starting point | Why |
|---|---|---|
| One brand, one principal market and a small number of important buying questions | Controlled manual baseline | You need to learn what should be measured before automating it |
| Repeated monitoring across several products, competitors, markets or languages | Paid tracker | Collection, history and comparison are becoming an ongoing operational burden |
| The company does not agree on the buyer questions, offer, evidence or meaning of success | Scoped independent audit | A subscription will automate an unresolved measurement and ownership problem |
| A dashboard already exists but nobody trusts or acts on it | Method and evidence review | The problem may be definitions, sampling or missing raw evidence rather than data volume |
Do not choose only by subscription price. Compare it with the time needed to run the protocol, check evidence, explain changes and coordinate corrections. A managed service should likewise show what was tested, observed and left uncertain - not hide the method behind a score.
5. Seven questions to ask before buying
Use this evidence test during a demo or proposal review.
| Question | Evidence worth requesting | Warning sign |
|---|---|---|
| Where do the questions come from? | Buyer, search, sales or research provenance; branded balance; market scope and exclusions | A large prompt count with no connection to your buying process |
| Which systems and conditions are tested? | Named products, interfaces, markets, languages, session conditions and dates | Results blended across undisclosed environments |
| How is answer variation handled? | Repeated observations for important questions, retained dates and visible ranges | One answer presented as a stable rank |
| What does the score count? | Definitions, formula, denominator and weighting | Mentions, recommendations and citations treated as equivalent |
| Can we inspect full answers and sources? | Raw responses, source URLs, exports and a change log | Only a score, chart or selected screenshot |
| Are accuracy and buyer fit reviewed? | Factual checks, recommendation criteria and documented valid exclusions | Every appearance counted as a win |
| What decision will this report change? | Named owner, correction path and review cadence | More monitoring with no operational next step |
A historical study of four generative search systems treated citation coverage and citation support as separate properties (Liu, Zhang and Liang, EMNLP 2023). Its results do not describe today’s market, but the distinction remains useful: a link is not automatic proof of accuracy.
6. What to fix after the measurement
Measurement earns its cost when it narrows the next decision.
If a page is not eligible or retrievable, inspect crawl access and indexability. If the company is described inaccurately, compare product pages, documentation, profiles and external sources. If competitors are recommended, inspect the criteria and sources supporting them before commissioning more generic content.
Some gaps cannot be solved by SEO alone. Product owns offer clarity, experts own evidence and communications influences external recognition. Someone must own the combined outcome. See GEO is not just an SEO problem.
If credible third parties used in answers do not know you, the work may involve research, partnerships, specialist communities or earned coverage - not schema. See what to build before AI visibility becomes more paid.
Do not confuse publishing an AI-facing file with proving retrieval or a useful outcome. Our llms.txt and WebMCP analysis explains the difference.
Correct the first explainable gap, keep the protocol stable and measure again. No movement is also information: the source may not have been retrieved, the evidence may remain weak or the correction may not affect the tested answer.
Conclusion: buy measurement after you define the decision
A small business does not need an enterprise dashboard to begin observing AI visibility. It needs real buying questions, clear definitions, stable conditions and evidence it can inspect.
Start manually. Learn which answers and sources matter. Buy software when repetition, scale and coordination justify automation. Commission an audit when the harder problem is deciding what to measure, why the gap exists and who can correct it.
The objective is not to own a better score. It is to make a better decision.
Establish an AI-visibility baseline your team can inspect
Bring us your principal market, buyer questions, current dashboard - if you have one - and the decisions the report needs to support. IZZY can scope a fixed AI Visibility Audit that separates eligibility, mentions, recommendations, citations and attributable visits, then prioritises the first explainable corrections.
No guaranteed citations and no invented universal rank: a bounded method, retained evidence and a clear next decision.
Frequently asked questions
Yes, once repeated monitoring across important questions, competitors or markets becomes operational work. Establish a manual baseline first so you know what is worth paying for.
Yes, within a limited scope. Test selected buyer questions under recorded conditions, inspect relevant first-party tools and measure attributable referrals. Call the result a bounded baseline, not total market visibility.
No. A platform can collect observations, identify cited sources and organise corrections, but it cannot guarantee a stable ChatGPT rank or recommendation.
A documented score can benchmark the questions, systems and conditions tested. It is not a universal measure of every AI conversation. Request the formula, question provenance, repeated results and raw answers.
There is no universal number. The set should represent the relevant buying decisions, markets, languages and constraints. The provider should explain its coverage, exclusions and fitness for the decision.
No. Search Console reports Google Search performance. ChatGPT visits should be inspected in analytics using available referral and UTM source information. OpenAI says its referral links include utm_source=chatgpt.com, but analytics still cannot observe unclicked mentions or private conversations.
Sources and product documentation checked on 20 July 2026. Product features can change. This article describes a measurement and purchase-decision framework; it does not guarantee a citation, recommendation, ranking or commercial outcome.