Pageoptimized
Module 11

AI visibility research and tracking

Design buyer-centered prompts, establish a repeatable testing protocol, and track mentions, accuracy, citations, competitors, and source gaps over time.

  • SEO specialists
  • Brand teams
  • Researchers
  • Agencies
Module case file

Horizon Legal wants to know why competitors appear in buyer answers

Five approved legal-service buyer questions are stored across 90 snapshots. The objective is to compare ChatGPT and Perplexity, inspect owned and external citations, and create work only where the evidence supports it.

Prompt set
5 approved Texas injury-law buyer questions
Competitors
Summit Injury Law and Northstar Trial Group
Tracked evidence
Response, mention, citation, cited page, accuracy
Cadence
Weekly; United States; fixed prompt text; saved raw evidence
Your finished deliverable

Finish a prompt protocol, diagnose one source gap, and create a history record that distinguishes movement from normal answer variation.

How this module works

The order matters

AI research discovers real buyer questions and source patterns. Prompt tracking observes a fixed, approved cohort over time. One exploratory answer is not a tracking baseline.

  1. 01Find real questions

    Cover discovery, comparison, constraints, risk, and post-purchase tasks.

  2. 02Control the test

    Fix prompt text and record engine, mode, date, location, and account state.

  3. 03Inspect the evidence

    Compare responses, sources, competitors, owned coverage, and factual consistency.

  4. 04Repeat and review

    Measure stable cohorts while keeping visibility separate from revenue.

Lesson 11.1

Research real buyer questions

Build prompts from audience decisions, not a list of brand-friendly formulations.

The prompt set decides everything downstream. Track the wrong questions and every number after them is precise, comparable, and about nothing anybody cares about.

The failure has a recognizable shape: a set built from what the company wishes people asked. Why is Horizon Legal the best Austin firm embeds its own conclusion, so the answer tells you nothing except whether an engine will repeat a premise you supplied.

What should I compare before hiring an Austin injury lawyer is the same territory and a completely different instrument. It reveals the decision criteria people actually use, and it shows you which competitors get named when nobody has tipped the scale.

Track the questions people ask, not the ones you would like answered. A prompt containing your conclusion cannot test anything.

Cover the decision, not the category

A useful set spans the stages of one real decision. Seven kinds of question, and most sets contain only the first and the last:

  • Discovery. Somebody with a problem and no vocabulary for it yet.
  • Comparison. Weighing named options against each other.
  • Constraint. A condition that changes the useful answer: budget, location, eligibility, timing.
  • Risk. What goes wrong, and what it costs when it does.
  • Implementation. How this actually works in practice.
  • Alternatives. What people choose instead, including doing nothing.
  • Post-purchase. What happens after the decision, which is where trust is won or lost.

Constraint prompts are the most informative and the least represented. An answer to what should I compare before hiring a lawyer is generic. Add I was rear-ended in Austin, have medical bills, and already have an insurer offer, and the answer becomes specific enough to be wrong in ways you can act on.

Group the set into cohorts by decision stage rather than by topic. A cohort is the unit you will compare over time, so it has to correspond to something you would act on separately.

Take the language from people, not from a tool

The words that belong in prompts come from recordings of sales calls, support tickets, community threads, and the queries in module 2's research. All of these are records of somebody trying to solve the problem in their own words, which is what a prompt is.

Remove anything written to flatter. Prompts that only exist to produce a favorable answer waste the run and corrupt the set. If a question would embarrass you to show a customer, it is not research, and the number it produces will be quoted internally as though it were.

Keep the set small enough to run repeatedly. Five approved prompts run weekly produce a comparable series; fifty run once produce a snapshot nobody can compare against anything. Repetition is worth more than coverage here, which is the reverse of module 2's keyword work.

One more filter is worth applying before a prompt is approved. Ask what you would do differently depending on the answer. A prompt whose result changes nothing is costing a run every week to produce a number that decorates a report, and a set assembled without that filter fills up with them faster than anybody expects.

Build the prompt set

Five steps. The output is a small approved set grouped into cohorts, not a long list.

  1. Collect the language people actually use.Sales calls, support tickets, community threads, and the qualified queries from module 2. Written language from real situations rather than category terms.
  2. Group questions by decision stage.Discovery through post-purchase. The gaps in your coverage become obvious the moment the stages are written across the top.
  3. Add the constraints that change the answer.Location, budget, eligibility, timing. These produce the specific answers worth diagnosing, and generic prompts produce generic responses about your whole category.
  4. Remove prompts that assume their answer.Anything containing your conclusion or your brand as the premise. They cannot fail, which means they cannot inform.
  5. Approve a small set and assign it to cohorts.Small enough to run on a schedule for months. The cohort is what you will compare, so it should map to a decision you would make separately.
Worked exampleSee the completed Horizon Legal work, then build your version.
Completed example: Horizon Legal

Buyer-question set

Questions come from real evaluation tasks and retain audience context.

Before this lesson: real audience language before prompts are written

Audience
Texas residents comparing injury counsel
Sources
Sales, support, search, reviews, and customer questions
Decision stages
Discover, compare, validate, and act
Boundary
No leading prompt written to force a brand mention

After this lesson: Finished output

Audience
Texas residents comparing injury counsel and evidence
Discover
Who are the best truck accident lawyers in Dallas for serious injury claims?
Compare
What should I compare before hiring a Houston car accident attorney?
Validate
Which Austin firms show proof of recent case outcomes?
Evidence
Which Texas injury lawyers publish settlement examples and client proof?
Decision

Approve five distinct buyer tasks before adding more prompts; avoid repetitive brand-friendly variants.

Save this in

AI Tracker: five approved Horizon Legal prompts with weekly history.

Your turn

Use the principle on your own project

Follow the sequence once. The goal is a defensible decision, not completing steps for their own sake.

Have these open

Sales, support, customer, community, search, and product language · Audience stages and decision tasks

  1. Collect customer, sales, support, community, and search language.
  2. Group questions by decision task.
  3. Include category, comparison, constraint, and branded questions.
  4. Remove prompts that only exist to flatter the brand.
Reference notesDefinitions, site-specific paths, common mistakes, and completion paths

Terms in plain language

Use these definitions when a term is unfamiliar.

Buyer question

A real decision, uncertainty, comparison, risk, or implementation need that can influence whether and how someone acts.

Example

What should I compare before hiring an injury lawyer after an Austin accident?

Prompt cohort

A deliberately grouped set of prompts representing one audience stage or decision, tracked together over time.

Example

Five prompts covering fees, case fit, alternatives, evidence, and first steps form an Austin provider-evaluation cohort.

Constraint prompt

A question including a condition that changes the useful answer, such as budget, location, risk, compatibility, eligibility, or timing.

Example

Which Austin injury lawyers handle truck accidents on contingency and offer Spanish-language intake?

Choose the path that matches your site

New sites establish evidence; established sites use history.

Brand-new site or no usable history

Use customer interviews, sales and support transcripts, community discussions, search language, competitor reviews, and the offer's real constraints. Label the first cohort as a hypothesis.

Site with usable history

Add converting queries, internal search, support cases, lost deals, content gaps, and current answer observations to select questions that affect decisions.

Common mistakes

What people often do and what to do instead.

Writing only branded or leading prompts
InsteadInclude non-branded category, comparison, constraint, risk, implementation, and alternative questions.
Tracking every brainstormed question
InsteadResearch broadly, then approve only distinct prompts tied to a decision.

You should now have

  • Prompt groups by decision task
  • Leading prompts removed
  • A candidate set ready for protocol review

Before you move on, confirm

  • Prompts map to real decisions.
  • The set includes non-branded questions.
  • Leading or manipulative wording is removed.
Lesson 11.2

Define the testing protocol

Make repeated observations comparable.

Module 10 established that one answer is an observation under conditions. This lesson is about writing those conditions down before you start, so that two runs six weeks apart are measuring the same thing.

Without a protocol, every difference between runs has at least three explanations and no way to choose between them. The engine changed, the conditions changed, or the world changed, and a report that cannot separate those is describing its own inconsistency.

Write the conditions before the first run. A difference you cannot attribute to a controlled variable is not a finding.

Record the conditions, then the result

Engine, product mode, account state, location, date, and the exact approved prompt text. These are cheap to record at run time and impossible to reconstruct afterward, which is why they get skipped and why skipping them is expensive.

Account state is the one people forget. A signed-in account with history produces different answers from a clean session, and a team that runs tests from whichever browser is open is mixing two populations without knowing it. Decide which one you are measuring and hold it.

Repetition is what makes it evidence

These systems return different answers to identical prompts. Running each prompt once and comparing weeks produces a chart of variance with a narrative attached to it.

Horizon runs five approved prompts weekly in the same products and country, and the tracker holds ninety snapshots across them. That volume is what turns a mention into a mention rate, and a rate is the smallest unit here that supports any claim at all.

Set the cadence to the decision, not to the dashboard. Weekly is right for an active program. Monthly is fine for a stable one. Daily produces cost and noise in equal measure, and nobody acts on a daily AI visibility number.

Version breaks are findings, not gaps

Products change. A search mode is added, a model is upgraded, a default changes, and the answers shift for reasons that have nothing to do with your site.

When that happens, record a protocol version break at that date and treat the series as two segments. The alternative is what usually happens: the shift lands in the middle of a campaign, gets reported as campaign impact, and the mistake is undiscoverable a quarter later because nobody wrote down what changed.

Define the classification rules in advance

Decide before the first run what counts as a mention, a recommendation, a citation, and a competitor appearance. Written afterward, these definitions bend toward whatever the data happened to show.

The edge cases are where the definition earns its keep. Does a mention of your parent company count? Does a citation of a third-party page describing you count as your citation? There is no universally right answer, and any consistent answer is better than deciding case by case for eight months.

Preserve every raw response and its cited sources alongside the classification. The classification is a judgment somebody made; the raw text is the evidence, and only the second one survives a disagreement.

Write the protocol

Six decisions recorded once, then held. Changing any of them starts a new series rather than continuing the old one.

  1. Fix the engines and product modes.Named specifically, because a product's search mode and its default mode are different instruments returning different answers.
  2. Fix the account state and location.Signed in or clean, and which market. Both change answers materially and neither is visible in the output.
  3. Approve the prompt text and lock it.Character for character. An edited prompt is a new prompt, and comparing it to its predecessor is comparing two questions.
  4. Set repetition and cadence.Enough runs per prompt to produce a rate rather than an anecdote, on a schedule matched to how fast you would act.
  5. Write the classification rules, including edge cases.Parent companies, third-party pages, partial matches. Decide once, in advance, and apply it consistently even where it is unflattering.
  6. Preserve raw responses and sources.Every run, verbatim. The saved text is what lets somebody re-derive your numbers, and re-derivable numbers are the only kind worth reporting.

Freeze a comparable AI Tracker protocol

Approve exact prompt text and record the environment fields that can change the answer. Edit the prompt only by creating a new version and baseline.

Record these fields
  • Prompt -> exact approved text
  • Engine and mode -> named for every run
  • Location and account/session state -> recorded
  • Cadence -> Weekly; raw response and sources retained
  • Classifications -> Mention / Citation / Accuracy / Recommendation
Do not compare these as one trend
  • A paraphrased prompt each week
  • Different engines combined without segmentation
  • One visibility score without the raw responses and sources
Worked exampleSee the completed Horizon Legal work, then build your version.
Completed example: Horizon Legal

Reproducible AI visibility protocol

The protocol keeps changing answer conditions visible.

Before this lesson: the candidate questions from Lesson 11.1

Candidate set
Five distinct buyer tasks
Variables
Prompt, engine, mode, location, session, and date
Required evidence
Raw answer, citations, cited URLs, and accuracy
Decision
Which fields stay fixed across runs?

After this lesson: Finished output

Cadence
Weekly
Prompt
Exact approved text; no silent edits
Environment
Engine, model/mode, location, session state recorded
Capture
Full response, mentions, citations, cited URLs, accuracy
Current set
5 prompts, 90 snapshots, latest Aug 28, United States
Decision

Compare only runs with matching protocol fields; label platform variation explicitly.

Save this in

AI Tracker schedule and Prompt Lab version history.

Your turn

Use the principle on your own project

Follow the sequence once. The goal is a defensible decision, not completing steps for their own sake.

Have these open

Approved prompt text and segments · Engine, mode, account state, location, date, cadence, repetitions, and classification rules

  1. Record engine, product mode, account state, location, and date.
  2. Use approved prompt text and segments.
  3. Choose repetition and cadence.
  4. Define mention, recommendation, citation, accuracy, and competitor rules.
  5. Preserve raw responses and sources.
Reference notesDefinitions, site-specific paths, common mistakes, and completion paths

Terms in plain language

Use these definitions when a term is unfamiliar.

Protocol

The recorded rules for running a test so later observations are as comparable as the product allows.

Example

Fixed prompt text, engine, product mode, country, account state, cadence, date, response, sources, and accuracy review.

Snapshot

One saved response and its metadata at one point in time. It is an observation, not a trend.

Example

Prompt P-04 in Perplexity on Aug 28 with response text and three cited sources is one snapshot.

Volatility

Normal variation in responses, sources, ordering, or wording between otherwise similar tests.

Example

A brand appears in four of eight weekly snapshots and changes order, indicating unstable visibility.

Choose the path that matches your site

New sites establish evidence; established sites use history.

Brand-new site or no usable history

Create the protocol and collect initial snapshots before major work ships. A zero-mention baseline is valid when prompt relevance and source review are still recorded.

Site with usable history

Reuse fixed cohorts, preserve prior snapshots and product-mode changes, and annotate releases so movement can be compared without erasing normal variation.

Common mistakes

What people often do and what to do instead.

Changing prompts between runs without versioning
InsteadFreeze approved text or create a new version and baseline.
Saving only summary metrics
InsteadPreserve raw responses, sources, timestamps, and run context.

You should now have

  • A versioned prompt cohort
  • Definitions and cadence
  • Raw-evidence retention and human review rules

Before you move on, confirm

  • The protocol can be repeated.
  • Metrics have definitions.
  • Raw evidence is retained.
Lesson 11.3

Diagnose source and answer gaps

Turn missing or inaccurate visibility into testable hypotheses.

The diagram above separates five reasons a brand is absent from an answer. This lesson is about the distinction underneath them, which decides where the work goes: whether you have a source gap or an answer gap.

An answer gap is about the response. You are missing, misdescribed, or framed badly. A source gap is about the evidence the engine had available: nothing useful existed, or what existed contradicted itself, or somebody else's page was simply better material.

A source gap is fixed by evidence existing somewhere. An answer gap may not be yours to fix at all.

Read the citations before the answer

The cited sources tell you more than the text does. Horizon's tracker shows three cited sites across the latest runs, of which one is theirs and two are other sites. That single line is the diagnosis: the engine is finding material on this topic, and mostly not from them.

Look at what the cited pages actually do. For the fee prompts, competitors' reviewed fee explanations are cited repeatedly and Horizon has no dedicated fee page at all. That is a coverage gap with an obvious response, and it is a response worth making because clients ask about fees constantly, not because publishing forces a citation.

The opposite pattern needs a different response. When your page is cited and the answer still describes you wrongly, the source is being read and misused, or the source itself is unclear. That is an editing problem on a page you own, which is the cheapest fix available in this module.

The mental model

Absent is not one diagnosis

What you can seeThe brand is not mentioned in the answer.
The prompt is not about you
Only this cause looks like

No competitor from your category appears either.

So you

Fix the prompt set. Nothing on the site is wrong.

Coverage gap
Only this cause looks like

The cited sources answer a question your site has never covered.

So you

Write that page. Only that page.

Source gap
Only this cause looks like

Competitors are named through third-party pages you are absent from.

So you

Earn a place in the source that gets cited, not another page you own.

Contradiction
Only this cause looks like

Your own pages disagree with each other on scope, price, or availability.

So you

Fix the fact first, then retest. Accuracy and mention are separate metrics.

Instability
Only this cause looks like

Same prompt, same day, different answers across runs.

So you

Repeat under the protocol before calling anything a gap at all.

Five different problems produce the same empty answer, and four of them are not fixed by writing another page.

Choose the smallest verification step

Every diagnosis here is a hypothesis, and a hypothesis predicts what should change if it is right. Before commissioning work, ask what the cheapest thing is that would confirm or kill it.

For a suspected instability, the cheapest test is running the same prompt three more times today. For a suspected coverage gap, it is checking whether the engine cites anybody at all for that question. Both cost minutes and both routinely prevent a month of content work aimed at the wrong cause.

Publish because the audience needs it. The reviewed fee resource is worth building on its own merits. Framing it as a citation play sets an expectation nobody can guarantee, and it will be judged against that expectation rather than against whether it helped a single client understand what they would pay.

Actual PageOptimized screenUse the numbered steps with the product image below.
Actual product · four evidence views

Inspect the sources behind visibility

The Citations view shows which domains and pages support answers, where competitors appear, and which source gaps deserve investigation.

PageOptimized AI Tracker citations view showing cited domains, pages, source coverage, and competitor evidence.Open full size
  1. Choose a prompt cohort.
  2. Open cited sources.
  3. Verify what each source supports.
  4. Create a hypothesis with a retest date.

The gap that is not yours

Some absences are correct. A prompt asking who handles class actions in Delaware should not name an Austin personal injury firm, and an engine that leaves you out of it is working. Every prompt set contains a few of these and they quietly drag the mention rate down without indicating anything.

The signal that separates them is whether anybody comparable is named. When the answer lists four firms and none of them are you, that is a gap. When it lists nobody from your category at all, the prompt is measuring something else and belongs out of the cohort.

Third-party sources are a different remedy. Competitors named through directories, review sites, or press coverage you are absent from is not a content problem on your site, and writing another page will not address it. That is module 7's work, and misdiagnosing it as a coverage gap is how a team publishes six pages and moves nothing.

Diagnose one absence

Four steps, and the first two are reading rather than doing. Most absences are diagnosed without touching the site.

  1. Read the full response and its cited sources.Not the classification. The sources name what the engine considered adequate, which is the most direct evidence available about what is missing.
  2. Compare competitor mentions and their source domains.Whether they are cited through their own pages or through third parties. Those are two different problems with two different remedies, and module 7 owns the second.
  3. Check your own coverage and consistency.Whether a page exists at all, and whether your pages agree with each other. Contradiction between your own pages is a common and entirely self-inflicted cause.
  4. Pick the smallest step that would test the hypothesis.Usually a rerun or a check rather than a content project. Confidence should be established before spend, not after it.
Worked exampleSee the completed Horizon Legal work, then build your version.
Completed example: Horizon Legal

Engine-gap diagnosis

The same prompt has different outcomes across engines, so the evidence does not justify a generic content-gap claim.

Before this lesson: a protocol-controlled cross-engine difference

Prompt
What should I compare before hiring a Houston attorney?
ChatGPT
Horizon named and cited
Perplexity
Horizon neither named nor cited
Need
Inspect responses and sources before prescribing work

After this lesson: Finished output

Prompt
What should I compare before hiring a Houston car accident attorney?
Observed
ChatGPT names and cites Horizon; Perplexity does not name or cite Horizon
Cited sources
Across the set: horizonlegal.com, avvo.com, and summitinjurylaw.com
Evidence shape
Named on every engine for 4 prompts; named on some for 1
Confidence
Medium; source gap observed, selection cause unknown
Decision

Inspect the two raw answers and recurring sources before changing the Houston page. Do not create a duplicate page only to chase one engine.

Save this in

AI Tracker Diagnosis & gaps -> verification task with the prompt, engine, response, sources, and retest date.

Your turn

Use the principle on your own project

Follow the sequence once. The goal is a defensible decision, not completing steps for their own sake.

Have these open

Full response and cited sources · Competitor mentions, owned pages, brand facts, and repeated-run stability

  1. Inspect the full response and cited sources.
  2. Compare competitor mentions and source domains.
  3. Check owned-page coverage and factual consistency.
  4. List plausible explanations.
  5. Choose the smallest verification or improvement step.
Reference notesDefinitions, site-specific paths, common mistakes, and completion paths

Terms in plain language

Use these definitions when a term is unfamiliar.

Source gap

Missing, weak, contradictory, inaccessible, or less useful evidence among the sources an answer product retrieves or cites.

Example

Competitor fee explanations are cited while Horizon has no clear, reviewed fee page available to retrieve.

Answer gap

A missing, inaccurate, poorly framed, or irrelevant treatment of the brand or topic in the observed response.

Example

The brand is mentioned but its Austin availability is stated incorrectly.

Hypothesis

A testable explanation that predicts what evidence should change if it is correct.

Example

If missing fee evidence is the gap, publishing a reviewed fee explanation should first improve source availability; citation change is then observed separately.

Choose the path that matches your site

New sites establish evidence; established sites use history.

Brand-new site or no usable history

Expect many source and awareness gaps. Prioritize questions central to the offer, publish accurate owned evidence, establish real independent references, and avoid manufacturing citations.

Site with usable history

Compare prompts, cited sources, current pages, source access, claim consistency, competitors, organic queries, and answer accuracy before choosing content or outreach.

Common mistakes

What people often do and what to do instead.

Calling every absence a trust or authority problem
InsteadCheck prompt relevance, source availability, factual consistency, retrieval variation, and instability.
Creating content from a single missing mention
InsteadChoose the smallest verification step and retest before committing broad work.

You should now have

  • Plausible explanations ranked by evidence
  • Owned and third-party source gaps
  • A proportional verification or improvement task

Before you move on, confirm

  • Absence is not automatically called a trust gap.
  • Competitor evidence is inspected.
  • The action is proportional to confidence.
Lesson 11.4

Track change and accuracy

Measure repeated observations beside the work and organic search evidence.

Once a protocol is running, the tempting move is to reduce it to one number and put that number in a report. This lesson is about which numbers are real, what they can be compared against, and the specific way a single score goes wrong.

The short version: track several rates separately, keep them away from rankings and revenue, and never let an unmeasurable component quietly count as zero.

Rates move independently and mean different things. A rising mention rate with flat citations and a false claim is not a win.

Four rates, tracked apart

Each of these is a share of recorded snapshots in a defined cohort, and each answers a different question:

  • Mention rate. How often you are named. Says nothing about sentiment or accuracy.
  • Citation rate. How often a visible source reference points at your site.
  • Accuracy. Whether what was said is correct, current, and appropriately framed. This one needs a person.
  • Competitor presence. Who else appears, which is what turns your rate into a position.

Horizon's case is the reason to keep them apart. Mentions rise from 2 in 10 to 7 in 10 across three consecutive weekly runs, citations stay at 1 in 10, and one claim is wrong. Reported as a single visibility number that reads as a large win, and it is at best a mixed result with a correction task attached.

Citation rate is the one worth watching hardest, because it is the only one of the four that depends on your own pages being good enough to reference. A mention can come from anywhere, including a competitor comparison that names you as the weaker option. A citation means something you published was worth pointing at.

What a composite score may and may not do

A blended score is defensible when its components are published and each one is genuinely measured. Named, cited, named first, and share of voice against tracked competitors, per engine, with the weights written down.

An unmeasurable component must be dropped, never zeroed. If sentiment cannot be assessed for an engine this period, the score for that engine is computed without it and says so. Treating the missing component as zero manufactures a decline out of an absence of data, and that decline will be investigated as though it were real.

Publish the weights wherever the score appears. A composite whose formula is not visible cannot be argued with, which sounds convenient and means nobody can tell you when it has started measuring the wrong thing.

Actual PageOptimized screenUse the numbered steps with the product image below.
Product source · AI Tracker

Read the cited sites, not just the mention count

The overview separates what is actually being measured: five approved prompts, how many were visible in the latest runs, how many sites were cited and whose they were, competitor hits across the tracked set, and the snapshot count behind all of it. Referral traffic states that analytics is connected but nothing has synced yet, rather than drawing a zero.

PageOptimized AI Tracker overview showing five tracked prompts, five visible prompts, three cited sites of which one is the project's own pages, eight competitor hits across two tracked competitors, ninety snapshots, and a referral traffic panel stating that no traffic sources have synced yet.Open full size
  1. Read cited sites as one of yours against two others.
  2. Check the snapshot count before trusting any rate.
  3. Keep competitor hits beside your own visibility.
  4. Treat an unsynced panel as unknown rather than as zero.

Keep it separate from rankings and revenue

AI visibility is its own layer. It is not organic ranking, and connecting it to revenue requires exactly the evidence module 1 demanded of every other connection.

The honest presentation puts it beside the other layers rather than inside them, with its own protocol, its own cohort, and its own dates. Horizon's tracker makes this concrete: referral traffic by AI engine sits on the same screen, and when analytics is connected but no traffic has synced, it says so instead of drawing a zero.

Actual PageOptimized screenUse the numbered steps with the product image below.
Actual product

Move from metric to diagnosis

Diagnosis panels keep engine gaps, losses, source patterns, and recommended verification beside the tracked evidence.

PageOptimized AI Tracker diagnosis view showing engine gaps, prompt losses, source patterns, and follow-up evidence.Open full size
  1. Find the changed cohort.
  2. Review the raw evidence.
  3. List plausible causes.
  4. Assign a verification task.

Attach the work and report the uncertainty

Record what shipped during each period alongside the movement, with dates. Not as proof, but so the plausible explanations are visible when somebody asks what caused a change.

Then state the confidence and the next verification date. A number reported without either is asking to be treated as more solid than it is, and in this part of the discipline that is the specific way credibility gets spent.

Report a tracking period

Six steps. The output should be legible to somebody who has never opened the tracker.

  1. Confirm the cohort and protocol were stable.Same prompts, same conditions, no version break. If any changed, the period is two segments and should be reported as such.
  2. Report the rates separately.Mention, citation, accuracy, competitor presence. A composite may accompany them and should never replace them.
  3. Have a person review accuracy.On the material claims. This is the one that cannot be automated and the one whose failure costs the most.
  4. Compute any score with its components visible.Published weights, and unmeasurable components dropped rather than zeroed. A score that hides a missing input reports a decline that did not happen.
  5. List the work shipped in the period.With dates, as context rather than as attribution. Module 1's join labels apply here without modification.
  6. State confidence and the next verification date.Including the alternative explanation you could not rule out. Then actually run the verification, which is what separates a program from a report.
Worked exampleSee the completed Horizon Legal work, then build your version.
Completed example: Horizon Legal

AI visibility change record

Movement and factual accuracy are reviewed separately.

Before this lesson: the approved baseline and diagnosis rules

Cohort
5 prompts; United States; weekly
History
90 saved snapshots
Measures
Mention, citation, source, competitor, and accuracy
Boundary
Visibility is not revenue and variation is expected

After this lesson: Finished output

Prompt set
Texas injury-law evaluation; 5 prompts
Visibility
5 of 5 visible; score 76; 90 saved snapshots
Citations
10 owned citations and 5 external citations in the latest set
Accuracy
No completed human accuracy verdict is stored for this example
Protocol
Same engine, mode, prompts, and weekly cadence
Decision

Report the current cross-engine pattern, not a gain. Add an accuracy task only after a reviewer identifies a specific incorrect claim.

Save this in

AI Tracker weekly history plus a human-review task when a claim needs verification.

Your turn

Use the principle on your own project

Follow the sequence once. The goal is a defensible decision, not completing steps for their own sake.

Have these open

Approved protocol and baseline · Human accuracy review, source records, competitor set, and shipped-work timeline

  1. Approve the tracked set.
  2. Run on a defined cadence.
  3. Review mention, accuracy, citation, sentiment, and competitor movement separately.
  4. Attach work shipped during the period.
  5. Report uncertainty and next verification.
Reference notesDefinitions, site-specific paths, common mistakes, and completion paths

Terms in plain language

Use these definitions when a term is unfamiliar.

Mention rate

The share of recorded snapshots in a defined prompt cohort where the brand is named. It does not state sentiment, recommendation, or citation.

Example

Horizon appears in 6 of 10 weekly comparison snapshots: a 60% observed mention rate for that cohort.

Citation rate

The share of snapshots containing a visible source reference to the tracked site or page under the recorded protocol.

Example

Two of ten snapshots cite horizonlegal.com even though six mention the brand.

Accuracy review

A human check of whether material claims in a response and source attribution are correct, current, and appropriately qualified.

Example

The answer names the correct office but states an unsupported case result, so mention is positive while accuracy fails.

Choose the path that matches your site

New sites establish evidence; established sites use history.

Brand-new site or no usable history

Collect enough fixed-protocol snapshots to understand normal variation before setting change expectations. Track zeros, source emergence, and accuracy without inventing benchmarks.

Site with usable history

Compare cohorts and protocol versions beside shipped work, organic evidence, source changes, and accuracy. Keep answer visibility separate from rankings and business outcomes.

Common mistakes

What people often do and what to do instead.

Reporting visibility as revenue
InsteadConnect it to business outcomes only when an attributable measurement source exists.
Combining mention, citation, accuracy, and sentiment into one score
InsteadReview each measure and its uncertainty separately.

You should now have

  • Comparable run history
  • Metric and source changes
  • Shipped work attached to the period
  • Next verification date

Before you move on, confirm

  • Cohorts are stable enough to compare.
  • Accuracy has a human review path.
  • Visibility is not reported as revenue without evidence.
Primary references

Verify the practice at the source.

Practices and source links reviewed August 2026.