In brief
Choose an AI visibility tool by how it collects answers, not by how its charts look. Ask which engines it covers and how, what each prompt costs, how often prompts repeat, whether it reads citations and competitors fairly, and whether it checks accuracy. The best tool turns answers into clear actions at a price you can see.
Key points
- How a tool collects answers, through APIs or consumer apps, with or without web search, shapes everything it reports.
- Cost per prompt per engine is the number that decides what you can afford to track.
- AI answers vary run to run, so a tool needs repeat runs to give a fair picture.
- Share of voice should weight position, since being named first is worth more than being named last.
- Tools that check whether answers are correct, not only whether you are named, catch a different class of problem.
- Look for actions and evidence you can act on, not only charts.
The steps at a glance
- 1
Write down what you need
List the engines, markets and number of prompts you need, and who will act on the findings.
- 2
Ask how answers are collected
Ask each vendor which engines they cover, whether through APIs or apps, and whether web search is on.
- 3
Work out cost per prompt
Ask for the cost of one prompt on one engine for one run, then multiply by your prompts, engines and frequency.
- 4
Trial on your own prompts
Run a short trial with your real buyer questions and competitors, and check the answers by hand.
- 5
Judge the actions
Look at what the tool tells you to do next, and whether it measures if those actions worked.
What does an AI visibility tool actually do?
An AI visibility tool asks AI engines the questions your buyers ask and records what comes back. It then reads each answer for your brand, your competitors and the sources cited.
That sounds simple. The differences between tools come from the details. Which engines are asked, how, how often, and what the tool does with the answers all change what you learn. This guide is a fair checklist for comparing tools, including ours.
People also ask whether GEO tools are worth buying or whether most are rebranded rank trackers. Some are close to rank trackers. A rank tracker records a position. A useful AI visibility tool reads a whole answer, which has no fixed positions, and tells you why you were or were not included.
Which engines does the tool cover, and how does it collect answers?
Start here, because it shapes everything else. Ask which engines are covered, and exactly how each one is asked.
There are two main ways to collect answers. One is through the engines' APIs. The other is by automating the consumer apps people use. APIs are stable and repeatable. App answers can include personal history, location and a different model version. Neither is perfect, so ask the vendor to explain their method and its limits.
Then ask whether web search is switched on. An engine with web search finds pages and usually cites them. An engine without it answers from what the model learnt in training. Both are useful, but they tell you different things. One shows whether your pages get found. The other shows how the model itself sees your brand.
Finally, check Google. AI Overviews appear on normal results pages. Ask whether the tool reads live Google results for your own keywords. Google AI Mode is harder, as Google offers no API for it, though its guide to generative AI features points to a Generative AI performance report in Search Console.
What does each prompt cost?
This is the number that decides what you can afford to track. Ask for the cost of one prompt, on one engine, for one run.
Many tools price by plan, with a set number of prompts. That is fine, but it can hide the real cost. Work it out by hand. Take the number of prompts you need, multiply by the engines, then by how often each runs in a month. Then compare that total with each plan.
A tool that publishes a price for every action makes this easy. A tool that cannot tell you what one prompt costs makes it hard to plan, and harder still to explain to a finance director.
How often are prompts repeated?
AI answers vary from run to run. The same question can name different brands on different days. A tool that runs each prompt once a month gives you a single snapshot, which may be unusual.
Ask how often each prompt runs and whether you can change the schedule. Weekly runs give a steadier picture than monthly ones. For your most important prompts, more frequent runs help you tell a real change from normal variation.
Be wary of any tool that presents one run as a definite result. Honest reporting treats tracking as a sample and shows trends over several runs. The guide on how to measure AI visibility explains why.
Does it track citations, not only mentions?
It should. Being named is one thing. Knowing which pages the engine relied on is what tells you what to do next.
Ask whether the tool records every cited page, not just whether yours was cited. Ask whether it groups sources by type, such as directories, review sites, listicles, social sites, news and competitor pages. That shows where your effort will count. If AI engines in your topic mostly cite two directories and a subreddit, that is where to look first.
How does it weight competitors and share of voice?
Ask how share of voice is worked out. A simple count treats a brand named first the same as a brand named last in a list of ten. That is not how buyers read answers.
A fairer method weights by position, so being named first counts more. Some tools also factor in whether a brand was recommended and how it was described. Ask to see the formula. If a vendor will not explain it, you cannot explain it to anyone else either.
Also ask how competitors are found. A tool that only tracks the competitors you enter will miss new ones. A tool that records every brand named in every answer will catch them.
Does it check whether answers are correct?
Most tools check whether you are named. Fewer check whether what is said about you is true. That is a different and often more urgent problem.
Wrong prices, closed branches, old services and mixed-up competitors all appear in AI answers. A tool that breaks answers into claims and checks them against facts you have approved can catch these. Ask how it decides what is wrong, how it avoids false alarms, and whether it helps find the page causing the error. The guide on when AI gets your business wrong covers why this matters.
Does it cover your locations and markets?
If you sell in more than one place, ask how the tool handles location. Can the same prompts run for different countries, regions or towns? Can you split results by market?
A business with eight showrooms needs to know it is named in one town and missing in the next. A national average hides that. Ask to see a report split by location during the trial.
Does it tell you what to do next?
Charts are easy. Knowing what to change is hard. The most useful tools turn findings into specific actions.
Look for things like pages to improve, questions you are missing, sources to target and facts to correct. Better still, look for tools that check later whether an action worked. Otherwise you have no way to tell good advice from busy work.
What should reports look like?
Reports need to make sense to someone who did not set up the tool. Ask to see a sample report before buying.
Check that it explains the method in plain words, shows trends rather than single runs, and separates engines. If you are an agency, check whether it can be white labelled and whether clients can log in to see only their own data.
What questions should you ask a vendor?
Use this table in sales calls. A good vendor will answer each one clearly.
| Question | Why it matters |
|---|---|
| Which engines do you cover, and through APIs or apps? | Method shapes every number you see |
| Is web search on for each engine? | Tells you whether answers reflect your pages or model knowledge |
| What does one prompt cost on each engine? | Lets you work out the real monthly cost |
| How often does each prompt run? | One run is a snapshot, repeats show trends |
| Do you record every cited page and classify it? | Shows where to act |
| How is share of voice calculated? | Position weighting gives a fairer picture |
| Do you check answers for accuracy? | Being named wrongly can be worse than not being named |
| Can I track by country, region or town? | National averages hide local gaps |
| What actions do you suggest, and do you measure results? | Separates insight from charts |
| Can I see the full price with no sales call? | Price transparency is a sign of how the vendor works |
Is a GEO tool worth it for your business?
For some businesses, a monthly manual check is enough to start. The AI visibility audit guide shows how to do one in a day. If that shows you are rarely named, or named wrongly, ongoing tracking starts to earn its cost.
A tool makes most sense when you need to track many prompts, several engines, several locations or several clients. It also makes sense when you need to show progress to a board or a client over months. The GEO research paper that named the field found visibility in generative engines can be changed by what you publish, and measurement is how you know whether yours has.
How Axiom GEO helps
Here are our own answers, briefly. Axiom GEO collects answers through the engines' APIs, with web search on for ChatGPT and Perplexity, model knowledge only for Gemini and Claude, and AI Overviews read from live Google results. AI visibility tracking runs your prompts on a schedule and records mention, position, sentiment, competitors, recommendation and cited pages. Competitor share of voice weights each mention by 1 divided by its rank. Answer accuracy, on the Professional plan and above, checks each claim against a fact sheet you approve. Every action has a published credit price at 1p a credit, so one prompt run once on all four engines costs 116 credits, and plan prices are on the pricing page.
Sources
These are the external pages this guide relies on.