AI Visibility Monitoring: A Reproducible Protocol
A fixed-prompt monitoring protocol for consulting firms that need dated evidence of AI mentions, citations, descriptions, and changes.
AI visibility monitoring is the repeated observation of a fixed, documented question set across named answer engines. Record whether the firm is named, how it is described, which sources are cited, and what changed between dated runs. Keep the prompt wording and run conditions visible beside each result. Do not invent a universal visibility score. A defensible record starts with raw captures under stable conditions, then adds summaries only when the underlying rows remain available. This article presents Jungle Roots' editorial method, not an industry standard.
What is AI visibility monitoring?
AI visibility monitoring asks a narrow question: what did a named answer engine return for a documented prompt under documented conditions on a documented date?
Use one captured answer as the unit of evidence. Give that answer its own row. The row says whether the firm appeared, whether a source was cited, how the firm was described, and when the capture was made. Do not turn an unobserved engine into a zero or treat one answer as a permanent ranking.
This is narrower than ordinary search-performance reporting. Google's Search Console Performance report covers performance on Google Search and lets site owners examine metrics such as clicks, impressions, click-through rate, and average position (Google Search Console Help, Performance report). It should not be represented as a report of answers produced by third-party assistants. Google separately explains how its own AI features relate to websites and says that traffic from AI Overviews and AI Mode is included in Search Console's overall Web search reporting (Google Search Central, AI features and your website).
If you need the wider audit frame before starting a recurring log, read the AI visibility audit. For the distinction between being named and being accurately represented, see brand visibility in AI search.
What should you monitor?
Keep the following observations separate:
- **Mention:** Did the answer name the firm or an unambiguous brand alias?
- **Citation:** Did the answer cite or link to a source associated with the answer?
- **Description:** What did the answer say the firm does, for whom, and in what context?
- **Change:** Which of those recorded fields differs from the previous comparable capture?
Do not collapse these observations too early. A mention without a citation is not the same event as a citation. A citation is not proof that the description is accurate. A changed answer is not yet proof of a durable change.
The record should also preserve the exact prompt, engine name, relevant condition, and capture date. These are not decorative metadata. In this method, they define the observation you actually made.
For a broader explanation of what can count as visibility across assistant answers, read LLM visibility.
How is AI visibility measured?
Start with rows, not a score. Use a measurement sheet with this structure:
| prompt | engine | condition | mention | citation | description | capture date |
|---|---|---|---|---|---|---|
| Exact frozen question | Named answer engine | Recorded access and session condition | Yes or no, with quoted evidence | URL, named source, or none observed | Short neutral transcription of how the firm is framed | Dated run |
| Exact frozen question | Named answer engine | Recorded access and session condition | Yes or no, with quoted evidence | URL, named source, or none observed | Short neutral transcription of how the firm is framed | Dated run |
The objection “one score proves visibility” fails an evidence check. A score hides the prompt, the engine, the condition, and the wording of the answer unless those inputs remain inspectable. Show raw captures and stable conditions before aggregation. If a summary is later useful, make it traceable to the rows and state exactly what was counted.
A first-party example shows why that restraint matters. The public Jungle Roots proof page displays a Perplexity baseline dated 2026-07-03 as 2/18. It also shows two cited brand-intent questions: “is Jungle Roots part of AI Jungle” and “what does Jungle Roots do for consulting firms.” That page explicitly makes no category-visibility extrapolation. Treat this as a dated, first-party baseline display, not as evidence of broader category visibility, a cross-engine result, or a customer outcome.
Which prompts belong in the baseline?
The baseline is the fixed question set against which later captures will be compared. Choose prompts from the firm's actual measurement question, then freeze their wording before the run.
For a professional-services firm, make the commercial context explicit. A frozen set could include:
- **Service discovery:** “Which cybersecurity consulting firms help mid-market healthcare providers prepare for NIS2 in France?”
- **Specialist comparison:** “How do specialist NIS2 consultancies for French healthcare providers differ?”
- **Buyer problem:** “Who can help a COO at a French healthcare provider close NIS2 readiness gaps?”
- **Branded due diligence:** “What does [Firm] do for healthcare providers preparing for NIS2 in France, and what proof supports that description?”
These are prompt examples, not observed results. For every prompt, record the service line (for example, NIS2 readiness advisory), buyer context (such as a COO at a mid-market healthcare provider), geography (France in this set), and the answer's exact description of the firm. Those fields make it possible to distinguish a missing mention from a visible but inaccurate position.
Baseline checklist
- Write each question exactly as it will be entered.
- Include the buyer context needed to interpret the answer.
- Separate discovery questions from questions that already contain the brand name.
- Keep category, comparison, problem, and brand questions distinguishable in the log.
- Assign each prompt a stable identifier.
- Record the answer engine beside each planned run.
- Define which access or session condition will be recorded.
- Freeze the baseline before interpreting results.
- Put newly discovered questions into a separate candidate list rather than rewriting prior rows.
This checklist is an editorial control, not a claim that every firm needs the same prompt categories. A consulting firm should be able to explain why every baseline question matters to the decision the monitoring is meant to support.
How often should runs be compared?
There is no universal frequency requirement in this protocol. Choose comparison dates based on the decision the evidence must support, then disclose the dates. The important constraint is comparability: rerun the fixed questions under conditions that are recorded in the same way.
Do not describe one snapshot as a trend. Do not compare a newly worded prompt with an old prompt and label the difference an engine change. When a material condition cannot be held stable, record the condition change and treat the row as non-comparable until you have evidence that supports another interpretation.
This approach avoids a false promise that a particular cadence makes the result representative. The log says what was observed and when. It does not claim that unobserved answers stayed unchanged between captures.
What should a capture record contain?
The measurement row is an index. The capture is the evidence behind it. Retain enough detail for another reviewer to understand the row without relying on memory.
A capture record should contain:
- the stable prompt identifier and exact prompt text;
- the named answer engine;
- the relevant access or session condition you observed;
- the capture date;
- the answer text or a faithful saved capture;
- the mention decision and the quoted passage supporting it;
- each cited source or link visible in the answer;
- a neutral description of how the firm was framed;
- the service line, buyer context, and geography represented by the prompt;
- a note when the run failed or the answer could not be captured.
Preserve “none observed” as a result. An empty citation field is ambiguous: it can mean no citation appeared, the reviewer forgot to check, or the capture is incomplete. Explicit language keeps those cases apart.
For Google Search's own AI features, Google says normal SEO best practices remain relevant and that there are no additional requirements or special optimizations needed to appear in AI Overviews or AI Mode (Google Search Central, AI features and your website). That guidance is about Google's AI features, so this protocol does not extend it into a claim about other answer engines.
How do you separate noise from change?
Call a difference “observed variation” before calling it a meaningful change. First check whether the prompt, engine, and recorded condition match the prior row. If they do not, add a condition-change note to the comparison.
Change-log checklist
- Link the new row to the prior comparable row.
- Confirm that the exact prompt text matches.
- Confirm that the named engine matches.
- Compare the recorded conditions.
- Mark whether the mention state changed.
- Mark whether the cited source set changed.
- Quote any material description change.
- Record the capture dates on both sides of the comparison.
- Label differences with changed conditions as non-comparable.
- Keep both raw captures available after writing the summary.
This is intentionally conservative. The log can prove that successive captured answers differed. Without additional evidence, it should not claim why the difference occurred or promise that the newer answer will persist.
When is a tool useful?
A tool is useful when it makes the fixed-prompt protocol easier to repeat without hiding the evidence. The tool should serve the method, not define the claim.
Tool due-diligence list
- Can you enter or import the exact prompt text without silent rewriting?
- Does each result retain the named engine and capture date?
- Can you record the access or session condition that matters to your run?
- Can you inspect the underlying answer rather than only a calculated score?
- Are mentions and citations stored as separate fields?
- Can you see the cited URLs or named sources attached to each answer?
- Can you preserve how the firm was described, including inaccurate wording?
- Can you export or otherwise retain prompt-level records?
- Does the tool disclose how any aggregate is calculated?
- Can you distinguish a failed capture from a genuine non-mention?
This list makes no claim about any vendor's capabilities, coverage, price, or reliability. Verify those details directly before adopting a product. A spreadsheet or manual log can also implement the protocol if it preserves the required fields; this is a workflow option, not a comparative tool claim.
What should an audit hand off?
An audit should hand off evidence that can become a monitoring baseline, not just a presentation. Include the frozen prompt set, completed measurement rows, raw captures, condition notes, and a change log ready for the next comparable run.
It should also separate observations from actions. “The firm was not cited in this captured answer” belongs in the evidence. “Review the page that should support this question” belongs in an action list. The separation prevents a proposed fix from being mistaken for a proven cause.
Map that action list to the firm's commercial materials. Route service-discovery and buyer-problem questions to the relevant service pages; route claims that need substantiation to named case studies, credentials, testimonials, or other proof assets; and flag descriptions that misstate the service line, buyer, or geography as positioning risks. The handoff should name the affected page or asset and preserve the inaccurate wording, without claiming that an edit will change a future answer.
For work on the source side of that action list, read how to get cited by ChatGPT. Treat that article as guidance, not as a guarantee that an edit will produce a citation.
A useful handoff answers these questions without requiring the original reviewer:
- What exact questions were asked?
- Where and under what recorded conditions were they asked?
- What did each captured answer say?
- Which sources were cited?
- Which rows can be compared later?
- What remains unknown?
FAQ
Is AI visibility monitoring the same as rank tracking?
Not under this protocol. The record follows captured answers to fixed questions and separates mention, citation, description, conditions, and date. It does not assume that every answer engine exposes one stable rank.
Can Google Search Console monitor third-party assistant citations?
Google documents the Search Console Performance report as reporting Google Search performance (Google Search Console Help, Performance report). Google also says traffic from its AI Overviews and AI Mode is included in overall Search Console Web reporting (Google Search Central, AI features and your website). Neither source should be presented as documentation that Search Console captures answers or citations from third-party assistants.
Should mentions and citations be counted together?
Keep them separate in the raw record. A mention answers whether the firm appeared in the answer text, while a citation records the source the answer displayed. Any later aggregate should expose both underlying fields.
Does one AI visibility score prove that a firm is visible?
No universal score is claimed here. Ask to inspect the prompt set, named engines, recorded conditions, capture dates, raw answers, and aggregation rule. Without those inputs, the score cannot show which observations produced it.
What if the answer changes between runs?
Record the difference as observed variation, check prompt and condition comparability, and preserve both captures. Do not assign a cause or call the change durable without evidence.
What is the first step for a consulting firm?
Define the decision the monitoring must support, freeze a relevant question set, and create the measurement sheet before running it. If you need an audit structure first, use the AI visibility audit as the starting frame.
The output should be a reviewable evidence trail: fixed prompts, named engines, stable recorded conditions, raw captures, and dated comparisons. That is more defensible than a number whose inputs cannot be inspected.
*Written by Tileo, an operator who measures how AI assistants cite brands, on his own portfolio first.*
Related reading
