How to Run an AI Visibility Audit That Produces Decisions
A repeatable AI visibility audit method for consulting firms: record mentions, citations, sentiment, accuracy, and actions in a query-level ledger.
An AI visibility audit is a repeatable observation of how AI assistants represent your firm for the questions buyers ask. It records whether the firm appears, whether its website is cited, what the answer says, whether the facts are accurate, and which sources support the response. That scope follows the audit dimensions described by Ahrefs.
Do not reduce the result to one visibility score. A mention can be positive but uncited. A citation can support an inaccurate statement. An accurate answer can omit your firm entirely. Keep visibility, citation, sentiment, and accuracy as separate fields, then route each finding to a specific action.
Run the same saved queries in the same assistants under recorded conditions. ChatGPT search responses can include inline citations and a Sources panel, so both the cited URL and its role in the answer belong in the record (OpenAI). The output is a query-level ledger that another consultant can inspect and repeat.
What should an AI visibility audit answer?
For a boutique consulting firm, the useful question is not “Are we visible in AI?” It is “For which buyer situations do we appear, how are we framed, and what evidence should we improve?”
Ahrefs defines an AI visibility audit as a structured assessment of where a brand is mentioned, how often, how accurately, and from which sources across AI search platforms (Ahrefs). This article’s method turns that idea into an observation ledger. The ledger preserves the prompt, response, citations, and evaluation, rather than hiding them behind a dashboard score.
That distinction matters operationally. These four dimensions answer different questions:
- **Visibility:** Did the assistant name the firm in its answer? This article’s method records the observed mention as yes or no, without treating it as proof of market share.
- **Citation:** Did the assistant link to the firm’s domain, or only to third-party pages? This article’s method records every visible source URL and the claim it appears to support.
- **Sentiment:** Was the description favorable, neutral, mixed, or unfavorable? This is a reading of the response’s framing, not a visibility measure.
- **Accuracy:** Were checkable statements about services, locations, clients, pricing, or people correct against an approved source? This article’s method requires a reference URL or internal owner for every accuracy judgment.
You can use our broader guide to AI visibility to place the audit inside a measurement program. The audit itself is the evidence capture step.
How should you choose queries from buyer decisions?
Start with the decisions a prospect makes before contacting a consulting firm. Do not begin with every keyword in a rank tracker. The query set needs enough context to reveal whether an assistant understands the firm’s category, fit, evidence, and alternatives.
This article’s method uses five query families:
- **Category discovery:** “Which boutique firms help B2B software companies improve pricing?”
- **Problem discovery:** “Who can diagnose low conversion from enterprise demos?”
- **Comparison:** “Compare [firm] with [named alternative] for a founder-led engagement.”
- **Validation:** “What evidence supports [firm]’s claims about its work?”
- **Branded fact check:** “What services does [firm] offer, and who is it for?”
For each family, write the query as a buyer would ask it. Add the relevant market, company size, or constraint only when it changes the decision. Save the exact text. A rewritten prompt is a new observation, not another run of the old one.
Ahrefs recommends defining platforms, entities, and regions or languages before benchmarking (Ahrefs). Use that scope as the header of the audit. In this article’s method, every conclusion stays inside that defined geography, assistant, and prompt scope.
How do you use a prompt, query, and citation ledger?
In this article’s method, a screenshot is attached evidence, while the structured row is the working dataset. Store one row per response and keep the exact query, cited URLs, and raw capture together.
Audit ledger template
| Field | What to record | Example |
|---|---|---|
| Observation ID | Stable row identifier from this article’s method | AVA-2026-07-30-01 |
| Timestamp and timezone | When the response was observed | 2026-07-30 14:10 UTC |
| Assistant and mode | Product plus search or browsing mode shown in the interface | ChatGPT, search used |
| Account or session state | Signed in or signed out; new or continuing conversation | Signed in, new chat |
| Market and language | Location context and response language | United States, English |
| Query family | Category, problem, comparison, validation, or branded fact check | Comparison |
| Exact query | Verbatim user input | Compare Firm A and Firm B for pricing work |
| Run number | Repetition number under the same saved conditions | 2 |
| Brand mentioned | Yes or no | Yes |
| Position in answer | First mention and answer section | Second firm in list |
| Brand URL cited | Exact URL, or none | https://example.com/case-study |
| Other cited URLs | Every visible third-party source | https://publication.example/article |
| Citation-to-claim note | Which sentence or claim each source appears to support | Source shown beside capability claim |
| Sentiment | Favorable, neutral, mixed, or unfavorable | Neutral |
| Accuracy | Accurate, partly accurate, inaccurate, or unverifiable | Partly accurate |
| Accuracy reference | Approved page or owner used to check the statement | https://example.com/services |
| Raw evidence | Saved response text and screenshot path | captures/AVA-2026-07-30-01.md |
| Follow-up action | Owner, action, and review date | Update services page; content owner; review 2026-08-15 |
The example values are illustrative parts of this article’s method, not benchmark targets. Replace them with your own observations.
ChatGPT search may search the web, and search responses include inline citations plus a Sources button when those features are present (OpenAI). Record the URL exposed by the interface, not just a publication name. This article’s method also requires the sentence nearest the citation in the raw evidence.
For Google’s AI features, keep ordinary technical and content checks in scope. Google says the same foundational SEO practices apply to AI Overviews and AI Mode, with no additional technical requirements beyond eligibility for Google Search with a snippet (Google Search Central). Google also states that meeting its requirements does not guarantee crawling, indexing, or serving, so an action plan must not promise inclusion (Google Search Central).
If you want an operator to inspect the measurement system behind these observations, Get the operations audit.
How do you separate observation from interpretation?
Scorecards become misleading when unlike findings collapse into one number. This article’s method keeps four labels on every row and summarizes each label independently.
For visibility, count only explicit brand mentions. For citations, distinguish the firm’s domain from third-party domains. For sentiment, quote the words that justify the label. For accuracy, test each factual statement against a designated source of truth. If the approved source is missing or conflicting, mark the claim unverifiable and assign an owner. Do not guess.
Use this review sequence for each response:
- Highlight every explicit mention of the firm and its named services.
- List each cited URL and paste the nearby claim into the evidence file.
- Underline language that signals recommendation, caution, limitation, or criticism.
- Extract checkable facts, then compare them with the approved reference for that fact.
- Record omissions only when the query clearly called for the missing information.
- Assign an action only after the evidence and label are saved.
When reviewers disagree about tone, return to the saved passage and record the reason for the chosen label. Keep source problems separate from messaging problems. See how Jungle Roots presents proof: the underlying evidence remains visible.
How do you repeat observations without manufacturing certainty?
AI responses can differ across runs, and Google says its AI features may use different models and techniques, producing varying responses and links (Google Search Central). A single response is therefore an observation, not a stable property of the brand.
This article’s method uses three runs of each saved query during one audit window. Keep the assistant, mode, language, market context, and conversation state fixed. Store each run as its own ledger row. The number three is a protocol choice for this method, not a statistical standard or an estimate of response stability.
For ongoing review, Ahrefs suggests repeating an AI visibility audit monthly when capacity permits, or quarterly otherwise (Ahrefs). Pick one cadence, record it in the scope, and preserve the query set. In this article’s method, new queries belong to a separate cohort and are reported separately from the original query set.
Repeat an observation outside the regular cadence when you need to verify that a recorded fact has changed. Treat the new response as a fresh observation. Do not claim that a page edit caused a later mention or citation unless the evidence can isolate that cause.
How do you turn findings into actions?
Route each pattern to the team responsible for the underlying evidence. Ahrefs groups actions around content gaps, off-site mentions, misinformation fixes, and competitor analysis (Ahrefs). This article’s method uses the decision table below to make that routing explicit.
| Observed pattern | What it means in this article’s method | Next action | Owner |
|---|---|---|---|
| Mentioned, own site cited, facts accurate | Evidence is present for this query; no broader conclusion | Preserve the cited page and monitor the saved query | SEO or content |
| Mentioned, no own-site citation | The brand appears, but this response did not cite its domain | Inspect cited third-party sources; improve the relevant first-party evidence if it is thin | Content and digital PR |
| Not mentioned, competitors cited | A query-level visibility gap was observed | Compare cited pages with your relevant page; decide whether the topic matches the firm’s offer | Strategy and content |
| Mentioned, facts inaccurate | The response contains a factual risk | Correct the canonical first-party page; identify cited pages carrying the error; document outreach | Brand and subject owner |
| Favorable framing, weak evidence | Positive language is not backed by a visible proof source | Publish or strengthen a case study, credential, or named methodology that can be checked | Consulting lead and content |
| Unfavorable framing, accurate | The response reflects a real limitation | Decide whether to change the offer or explain the tradeoff clearly | Leadership |
| Unfavorable framing, inaccurate | The response presents a reputation and accuracy issue | Preserve evidence, correct owned sources, and request factual corrections from relevant publishers | Brand and communications |
| Conflicting results across runs | The observation is not consistent within this audit window | Keep all rows, report the split, and repeat at the next scheduled review | Analyst |
Technical work should follow the documented baseline, not a special “AI file” checklist. Google says no new machine-readable AI file or special schema is required for its AI features; crawl access, internal discoverability, textual content, page experience, and matching structured data remain among its stated SEO considerations (Google Search Central). That guidance supports good site maintenance, but it does not promise an AI Overview link.
Use collection tools only when they preserve the fields a reviewer needs. Our guide to LLM SEO tools explains how to assess tools without handing the decision to a single vendor metric.
How should you report the audit?
Lead the report with decisions, then attach the ledger. This article’s method uses a short findings page with four separate summaries: observed mention coverage, own-domain citation coverage, sentiment labels, and accuracy labels. Each summary links back to query-level rows.
Avoid a blended “AI visibility score.” If leadership requires a compact status view, use counts by label and show the denominator: observed rows in the defined audit scope. State the assistants, markets, query cohort, run protocol, and dates beside the counts. These are observations from a bounded test, not estimates of every answer a buyer might receive.
A useful finding reads like this: “In this audit window, the firm appeared in the recorded comparison-query rows, but its domain was not cited; the cited pages were two publisher comparisons.” It names the scope, evidence, and gap. It does not claim why the assistant produced the answer.
What are common AI visibility audit questions?
What is an AI visibility audit?
It is a structured check of where a brand appears in AI search responses, how it is described, whether its site is cited, whether statements are accurate, and which sources appear with the answer. Those dimensions align with Ahrefs’ audit definition (Ahrefs). This article’s method records them at query level so the findings can be inspected and repeated.
Which queries should a consulting firm test?
Use queries tied to category discovery, problem discovery, comparison, validation, and branded fact checks. Those five families are this article’s method, not a universal taxonomy. Include the market and buyer constraint when they are part of the actual decision, then freeze the exact wording for repeated runs.
How should citations be recorded?
Save every visible URL, the nearby claim, the response text, the timestamp, and a screenshot. For ChatGPT search, OpenAI documents inline citations and a Sources panel in search responses (OpenAI). Keep own-domain and third-party citations in separate fields.
How do you separate visibility, sentiment, and accuracy?
Use independent labels. Visibility asks whether the brand appeared. Sentiment describes the answer’s framing. Accuracy checks factual statements against an approved reference. A response can pass one test and fail another, which is why this article’s method does not blend them into one score.
How often should observations be repeated?
Within an audit window, this article’s method uses three saved runs per query under fixed conditions. For the audit cycle, Ahrefs recommends monthly reviews when capacity allows and quarterly reviews otherwise (Ahrefs). Keep the chosen cadence consistent and treat each response as an observation.
Does fixing a page guarantee more AI citations?
No. Google states that standard SEO fundamentals remain relevant to its AI features, while crawling, indexing, and serving are not guaranteed even when a page meets requirements and policies (Google Search Central). An audit identifies evidence gaps and assigns work; it does not prove that the work caused a later citation.
What actions should follow the audit?
Route inaccurate facts to the owner of the source-of-truth page, citation gaps to content or digital PR, technical issues to SEO or engineering, real offer limitations to leadership, and monitoring items to the analyst. Keep the triggering ledger rows attached to the task so the next review can compare like with like.
If you want the ledger, evidence trail, and operating decisions reviewed together, Get the operations audit.
*Written by Tileo, an operator who measures how AI assistants cite brands, on his own portfolio first.*
Related reading
