Pageoptimized
Module 10

AEO and GEO without myths

Improve accessibility, clarity, evidence, and source eligibility for search and AI experiences without inventing files, chunk formulas, or citation guarantees.

  • Business owners
  • SEO specialists
  • Content teams
  • Agencies
Module case file

Horizon Legal needs AI access and answer checks without folklore

The team wants visibility in answer engines. Site Health now preserves exact crawler-control tests and a failed-to-passed browser-task rerun, while AI Tracker holds separate prompt and citation observations.

Site
horizonlegal.com
Critical path
Service pages, case proof, and settlement calculator
Verified access
7 of 8 audited pages readable; calculator fails rendering
Answer evidence
5 prompts, 90 snapshots, named engines, responses, and sources
Your finished deliverable

Review the saved crawler matrix and browser-task rerun, then report only observed access, mentions, citations, and task completion.

How this module works

The order matters

AEO and GEO build on accessible, useful, evidence-backed pages. Crawler controls, agentic usability, and answer observations are separate checks; none guarantees a citation or recommendation.

  1. 01Protect the foundation

    Keep important information crawlable, rendered, useful, and current.

  2. 02Map controls

    Separate search, retrieval, training, robots, WAF, and rendering behavior.

  3. 03Test the task

    Verify that people and agents can perceive, operate, and recover from errors.

  4. 04Report honestly

    Record the protocol and distinguish observations from causal claims.

Lesson 10.1

Use the same strong foundation

Start with accessible, indexable, useful pages rather than an AI-only optimization checklist.

This is the module most likely to disagree with something you read last week. AEO and GEO have generated more confident advice per unit of evidence than any topic in search, and most of it consists of a specific tactic presented as necessary, with no product documentation behind it.

The uncomfortable answer is that the foundation has not changed. Answer engines still need to reach the page, read the content, and find something worth using. Accessible, indexable, genuinely useful pages remain the prerequisite, and almost every AI-specific tactic being sold is either a restatement of that or an invention.

What has changed is worth taking seriously, and it is covered properly in the next three lessons: different crawlers with different purposes, agent-driven interaction with your interface, and a category of observation that needs its own evidence discipline. None of it replaces the first seven modules.

There is no AI-only switch. Fix what makes the page usable and retrievable, and be suspicious of any tactic with no documentation behind it.

The tactics with nothing behind them

Four claims that circulate constantly and have no documented mechanism supporting them:

  • An llms.txt file is required. No major engine documents it as a requirement, and adding one changes nothing about whether you are retrieved.
  • There is a special AI schema type. Structured data works as module 5 described it, and no vocabulary exists that signals suitability for generated answers.
  • Paragraphs must be a prescribed length for chunking. Retrieval implementations differ, are undocumented, and change without notice.
  • Formatting alone earns citations. Adding a question heading to a paragraph does not make the paragraph worth quoting.

The common shape is a cheap, visible action offered as a substitute for expensive, invisible work. Each one is attractive precisely because it can be completed in an afternoon and cannot be disproved quickly.

None of these are harmful in themselves. The harm is the substitution: a team ships an llms.txt file, records the AI optimization task as complete, and never repairs the page that fails to render.

What actually helps, and why it is unglamorous

Horizon's highest-value AI work is repairing the settlement calculator so its explanation exists without scripts, and publishing the reviewed limitations of the estimate. Both are ordinary content and rendering fixes from modules 4 and 5.

They help because they change what is available to retrieve. A page that renders to nothing offers nothing to any system, generative or otherwise. A page carrying a reviewed, specific, attributable statement about Texas settlement limitations offers something quotable that a general article does not.

This is the whole mechanism, as far as anything is documented. Be reachable, be readable, and contain something worth using. Everything else on offer is a claim about internals that no vendor has published.

Actual PageOptimized screenUse the numbered steps with the product image below.
Free tool · AI-Readable HTML View

See the server HTML before you argue about AI access

It shows the title, headings, schema, links, and text present in the server HTML returned to PageOptimized's crawler. That is a transparent view of one fetch, not a claim about what any particular AI system does with your page, and the distinction is the whole lesson.

PageOptimized free AI-Readable HTML View tool showing a page URL input and a description covering the title, headings, schema, links, and text exposed in the server HTML returned to its crawler.Open full size
  1. Run it on a page that depends on scripts for its content.
  2. Compare what appears here with what a person sees in a browser.
  3. Treat a gap as a rendering finding from module 4.
  4. Do not read the output as an AI visibility verdict.

How to evaluate the next tactic you are told about

Three questions, in this order, applied to anything presented as necessary for AI visibility:

  • Which product documents this, and where? A vendor blog citing another vendor blog is not documentation.
  • What is the claimed mechanism, stated specifically enough to be wrong?
  • Would this help a human reader? If yes, do it for that reason and stop making the AI claim.

Most tactics fail the first question. Of those that survive it, most turn out to be ordinary good practice with a new name attached, which is fine as long as nobody is charging extra for the name.

The third question is the useful one to keep, because it dissolves most arguments without needing to settle the mechanism. Clear headings, plain answers early, and specific attributable facts are worth doing whether or not any engine rewards them, and a tactic that fails all three questions is asking you to spend effort on a claim nobody will ever be able to check.

Say eligibility, never guarantee

Nothing you do makes a citation certain. Engines choose sources per query, per model version, per mode, and the same page can be cited today and absent tomorrow with no change on your side.

That means the honest promise is about being retrievable and worth using, never about the outcome. A report that says the calculator now renders its explanation without scripts is a fact. A report that says this will get us cited by ChatGPT is a claim nobody can support, and it will be remembered when it does not happen.

Build the foundation

Five steps, none of which are new. That is the lesson rather than a shortcoming of it.

  1. Verify crawl and index controls.Module 4, unchanged. A page that cannot be fetched or that carries a stale noindex is not a candidate for anything, generative or otherwise.
  2. Make the primary information exist in rendered text.Not behind a click, a tab, or a script that might fail. This is the single highest-value AI-related fix on most sites and it is a rendering fix.
  3. Organize it for a person.Clear headings, one idea per section, plain language. Useful for readers, and it happens to be what makes a passage extractable.
  4. Support consequential claims with attributable evidence.Named sources, named reviewers, and dates. Specific and attributable content is what a system has a reason to quote rather than paraphrase from elsewhere.
  5. Keep important facts current.Stale facts propagate. Module 7's reputation work is the downstream cost of not doing this, and correcting a third party is far more expensive than being right first.
Worked exampleSee the completed Horizon Legal work, then build your version.
Completed example: Horizon Legal

AI discovery foundation check

The same project evidence is checked before inventing AEO-only work.

Before this lesson: critical pages before any AI-specific claim

Page set
Services, case proof, and settlement calculator
Verified access
7 of 8 audited pages readable
Known failure
Calculator content absent from initial HTML
Boundary
No llms.txt, special schema, or citation guarantee

After this lesson: Finished output

Critical pages
Service pages, case proof, and settlement calculator
Access
Seven pages expose readable content; /settlement-calculator/ is an unreadable JavaScript shell
Content
Case proof and service pages are cited in saved AI answers; accuracy still requires human review
Structure
Site Health stores titles, headings, internal-link counts, schema, and rendering state per URL
Decision

Fix weak or inaccessible source pages before adding any answer-engine tracking layer.

Save this in

Site Health and content QA record for the critical-page set.

Your turn

Use the principle on your own project

Follow the sequence once. The goal is a defensible decision, not completing steps for their own sake.

Have these open

Representative important pages · Crawl, render, index, content, source, and freshness evidence

  1. Verify crawl and index controls.
  2. Make the primary information available in rendered text.
  3. Use clear organization for people.
  4. Support consequential claims with attributable evidence.
  5. Keep important information current.
Reference notesDefinitions, site-specific paths, common mistakes, and completion paths

Terms in plain language

Use these definitions when a term is unfamiliar.

Answer engine

A system that generates or assembles an answer from model knowledge, retrieved sources, tools, and product-specific behavior.

Example

ChatGPT, Perplexity, Gemini, and Google AI features can use different retrieval and citation behavior.

Citation

A visible source reference attached to a particular answer observation. It does not prove endorsement, stable inclusion, or business impact.

Example

A Horizon Legal guide is cited for one prompt in one Perplexity snapshot.

AI-only optimization

A claimed tactic presented as necessary specifically for AI answers despite lacking reliable product documentation or evidence.

Example

There is no universal requirement to add llms.txt, special AI schema, or a fixed paragraph length for visibility.

Choose the path that matches your site

New sites establish evidence; established sites use history.

Brand-new site or no usable history

Prioritize accessible, indexable, useful pages, accurate identity, original evidence, and clear ownership. Establish a small prompt baseline only after real audience questions are known.

Site with usable history

Audit technical access, page usefulness, entity and claim consistency, organic search evidence, external sources, and repeatable answer observations before adding AI-specific work.

Common mistakes

What people often do and what to do instead.

Adding llms.txt, special AI schema, or fixed paragraph chunks as requirements
InsteadUse documented controls and improvements that help people and search systems.
Promising citations from formatting changes
InsteadTreat citation as an observed outcome under a repeatable protocol.

You should now have

  • A foundation readiness record
  • Documented gaps
  • No unsupported AI optimization claims

Before you move on, confirm

  • No undocumented requirement is presented as fact.
  • Changes improve the human experience too.
  • Claims and dates are reviewable.
Lesson 10.2

Build a crawler and control matrix

Distinguish search discovery, user-requested retrieval, training, and rendering controls.

One company can operate several crawlers doing unrelated jobs, and the names do not tell you which is which. Blocking the one you recognize while allowing the one that matters, or the reverse, is the standard outcome of guessing.

The remedy is a matrix you build from official documentation rather than from memory: which product, which robots token, which user agent, what it is for, whether it currently reaches you, and how you verified that.

Build the matrix from vendor documentation, and record how each access result was verified. Allowed in robots.txt proves almost nothing.

What allowed actually means

This is the most consequential misunderstanding in the module. A robots.txt test that comes back allowed tells you one thing: the token you tested was not disallowed by the rules you tested it against.

It does not tell you any of the following, each of which can independently prevent the outcome you wanted:

  • That the crawler actually requested the page.
  • That it got past the WAF or bot management, which frequently blocks what robots.txt permits.
  • That the response rendered into usable content.
  • That anything was stored or indexed.
  • That a citation was ever produced.

Five separate stages, and robots.txt governs the first gate of the first one. Module 1's staged thinking applies here exactly as it does to search.

So the access column in your matrix should record what you verified and how. Request logs showing the user agent arriving is evidence. A robots.txt checker returning green is a configuration statement, not an observation.

The mental model

One vendor, several crawlers, different jobs

Search discovery
What it is for

Finding and refreshing pages for a search index

What controls it

robots.txt and page-level index controls

What it does not do

Blocking it does not remove you from answers built from other sources

User-requested retrieval
What it is for

Fetching your page because someone asked a question right now

What controls it

Its own documented token, separate from search

What it does not do

Allowing it does not earn a citation

Model training collection
What it is for

Gathering text for future model training

What controls it

A separately documented training token

What it does not do

Blocking it has no bearing on how you rank

Agent and browser automation
What it is for

Completing a task in a browser on a person's behalf

What controls it

Authentication, consent, and confirmation boundaries in the product

What it does not do

It is not governed by ranking rules at all

Fill each row from the vendor's own documentation, and record the token, the user agent, and how you verified access. Names change; the purposes are what you are actually controlling.

Crawlers from the same company can serve unrelated purposes. Blocking one is not blocking the others, and none of it is a ranking decision.

Separate the controls, and check the WAF

Search discovery, user-requested retrieval, and training collection are governed by different tokens, and conflating them produces decisions nobody intended. A business that wanted to opt out of training and blocked everything with a similar name can remove itself from answer surfaces it wanted to be in.

The WAF is where most surprises live. Security tooling makes its own decisions about unfamiliar user agents, usually without telling anybody. A crawler permitted in robots.txt and refused by bot management looks identical to a crawler you blocked on purpose, right up until somebody reads the logs.

Recheck periodically. Vendors add crawlers, rename tokens, and split one product into two, and a matrix built eighteen months ago is describing a landscape that has moved. Put a review date on it like any other evidence.

Deciding what to allow is a business decision

The matrix describes the landscape. It does not tell you what to permit, and that question is commercial rather than technical, which is why it should not be settled quietly by whoever last edited robots.txt.

Blocking training collection has no documented effect on how you rank or whether you are retrieved for a live question. Organizations block it for reasons that have nothing to do with search: licensing positions, contractual obligations, or a view about their content being used to train a competitor's product. Those are legitimate and they belong to the business.

Blocking live retrieval is a different decision with a visible cost. That is the crawler fetching your page because somebody asked a question right now. Refusing it removes you from that surface while your competitors remain, and it is the choice most often made by accident, by somebody blocking a family of similarly named tokens in one edit.

Write the decision and the reason next to each row, and have somebody outside the search team agree to the ones about training and licensing. A rule with no recorded rationale gets reversed by the next person who reads the file and assumes it was a mistake.

Build the matrix

Six columns, filled from documentation rather than from what somebody recalls. Only list products that matter to this organization.

  1. List only the products relevant to you.A matrix covering every crawler in existence is unmaintainable and will be abandoned. Cover the surfaces this business actually cares about appearing in.
  2. Record the robots token from official documentation.From the vendor's own page, with a link. Tokens copied from a blog post are how a rule ends up matching nothing at all.
  3. Write the purpose in your own words.Search discovery, live retrieval, training, or something else. Writing it out is what exposes the cases where two similarly named crawlers do unrelated jobs.
  4. Test robots, status, WAF, and rendering separately.Four independent gates. A single pass or fail hides which one is actually deciding the outcome.
  5. Record the verification method beside each result.Logs, a published IP verification method, or a reproducible test. An access conclusion with no method behind it is an assumption in a table.
  6. Set a review date.The landscape changes without notice. A matrix nobody has revisited in a year is a historical document being read as current.

Use one row per crawler and purpose

Record only the verdict produced by that exact test. A robots rule, HTTP response, rendered result, and log observation are separate fields.

Required matrix fields
  • Product / crawler -> OAI-SearchBot
  • Purpose -> Search discovery, from official documentation
  • Controls -> robots token / WAF / HTTP / rendered content / logs
  • Verdict -> Observed result, date, method, and limitation
Invalid conclusions
  • GPTBot allowed -> ChatGPT search access proven
  • One AI answer -> crawler access proven
  • robots allowed -> WAF, rendering, and authentication passed
Worked exampleSee the completed Horizon Legal work, then build your version.
Completed example: Horizon Legal

Crawler control matrix

The saved Site Health matrix shows why crawler purpose and access layers cannot be collapsed into one allowed or blocked label.

Before this lesson: saved identity-specific access evidence

Verified
OAI-SearchBot readable; Claude-SearchBot WAF blocked
Control only
Google-Extended robots policy; no request UA
Controls
Robots, WAF, HTTP, render, logs, and purpose
Rule
Do not transfer one crawler's result to another

After this lesson: Finished output

OAI-SearchBot
Robots allowed; HTTP 200; WAF passed; 986 plain-HTML words; rendered result readable; one matching origin-log request
Claude-SearchBot
Robots allowed, but the live request returned HTTP 403 at the WAF; content therefore remains unavailable for this exact test
Google-Extended
Robots control blocked; no HTTP or rendering result because Google-Extended is not a request user agent
Evidence boundary
The documented user agent was simulated; provider IP identity, indexing, citation, and recommendation remain unproven
Decision

Keep crawler purpose, robots directive, HTTP result, rendered result, and review date in separate columns.

Save this in

Site Health -> AI access and agentic browsing; each rerun remains in the crawler's history.

Your turn

Use the principle on your own project

Follow the sequence once. The goal is a defensible decision, not completing steps for their own sake.

Have these open

Relevant search and AI products · Official crawler documentation, robots rules, WAF logs, status, and render tests

  1. List only products relevant to the organization.
  2. Use official crawler documentation.
  3. Separate search/discovery from training controls.
  4. Verify user agents and published IP methods where available.
  5. Test WAF, robots, status, and rendered access independently.
Reference notesDefinitions, site-specific paths, common mistakes, and completion paths

Terms in plain language

Use these definitions when a term is unfamiliar.

Crawler product

The named service or feature associated with a crawler. Similar company names do not mean identical purpose or controls.

Example

A search crawler, user-requested fetcher, and model-training crawler can have different documentation and robots tokens.

Robots token

The user-agent name documented for robots.txt matching. It must be recorded from official product documentation.

Example

Do not infer that one token controls every crawler operated by the same company.

Verification method

The evidence used to identify a crawler or access result, such as documented user agent, published IP method, request logs, response, and rendered output.

Example

A user-agent string alone can be spoofed, so stronger identity claims may require additional documented checks.

Choose the path that matches your site

New sites establish evidence; established sites use history.

Brand-new site or no usable history

Define intended search, retrieval, and training controls before launch, then test documented identities against production robots, WAF, status, and rendered responses.

Site with usable history

Inventory current rules, security behavior, server logs, and observed access. Test each product and purpose separately rather than copying one robots policy across all AI crawlers.

Common mistakes

What people often do and what to do instead.

Assuming every bot from one company has the same purpose
InsteadSeparate search discovery, user-requested retrieval, training, and other documented uses.
Testing robots.txt only
InsteadVerify DNS, WAF, status, authentication, rendering, and published IP methods when available.
Current workflow

Complete this in PageOptimized

What PageOptimized does
Site Health now saves one immutable test per named crawler identity, including official documentation, purpose, robots token and rule, HTTP and WAF result, plain and optional rendered content, limitations, reviewer, time, and optional log evidence.
What it does not do
A simulated documented user agent does not prove the request came from the provider's production IP range, and no access result proves indexing, training, citation, or recommendation.
How to use it now
Run each relevant identity against a representative URL, add server-log evidence when available, and report only the layer each field observed.

You should now have

  • A product-by-control matrix
  • Observed access and failure layer
  • Source links and retest date

Before you move on, confirm

  • Every crawler has a documented purpose.
  • Training controls are not described as search controls.
  • Access conclusions have logs or reproducible tests.
Lesson 10.3

Inspect AI view and test agent readiness

Check the static page signals that support automation, then preserve a real browser-task run with ordered evidence and a comparable rerun.

Software is starting to use interfaces on people's behalf: reading a page, filling a form, choosing from a control, completing or abandoning a task. Whether that works on your site has almost nothing to do with crawler permission and almost everything to do with how the interface is built.

The useful framing is that an agent is a user with no eyes, no intuition, and no patience for ambiguity. Everything that helps it is something that already helps a person using a screen reader or a keyboard, which is why this lesson is mostly accessibility work with a different motivation.

An agent is a user without eyes or intuition. What helps it is accessibility, not an AI-specific text block.

What actually decides whether a task can complete

Six properties of the interface, none of which are content and all of which decide the outcome:

  • Semantic structure. Headings and landmarks that describe the page rather than just styling it.
  • Accessible names on controls. A date picker whose buttons have no names cannot be operated reliably by anything.
  • Keyboard operability of the primary path, end to end.
  • Stable, announced state. What changed after an action, said in a way software can read.
  • Clear errors. What went wrong and what to do, next to the field that caused it.
  • Safe confirmation before anything irreversible, costly, or externally visible.

Horizon's failure is the fourth item's neighbor: the consultation date controls carry no accessible names, so a task told to find a slot cannot reliably choose a time. The fix is an interface change. No amount of text on the page addresses it.

Every one of these is already a requirement for people using assistive technology, which is the useful thing about this list. The work has an existing justification, an existing standard, and frequently an existing legal obligation, none of which depend on an argument about what any AI system does.

Test a real task, and record what happened

Pick the primary path a person would actually want automated, and give it a boundary. Find a consultation slot and stop before submitting is a good test task: it exercises navigation, controls, and state without taking an action on somebody's behalf.

Preserve the run, not the verdict. Every step result, the observed state, the semantic evidence, the screenshot, the point of failure, and what approval or execution source was involved. A pass or fail with nothing behind it cannot be compared against the rerun after the fix, which is the only thing that proves the fix worked.

Report failures as what they are: user-experience and automation-readiness problems. They are real, they affect real people using assistive technology today, and they do not need an unprovable ranking claim attached to justify the work.

Actual PageOptimized screenUse the numbered steps with the product image below.
Free tool · Agentic Page Readiness

Check the static signals an agent has to work with

The checker inspects static HTML signals that help a browser agent understand links, controls, forms, labels, and page structure. It is a preflight, and the tool says so itself: HTML readiness only, with no claim that a browser task succeeded.

PageOptimized free Agentic Page Readiness Checker showing a page URL input and a description covering static HTML signals for links, controls, forms, labels, and page structure.Open full size
  1. Run it on the primary task path, not the homepage.
  2. Check that controls carry accessible names.
  3. Treat a pass as a preflight rather than a completed task.
  4. Record a real browser-agent run in Site Health for the actual evidence.

Where an agent should be stopped

Not every task should complete, and a site that lets software do anything is not more ready than one that does not. Irreversible, costly, private, or externally visible actions need an explicit checkpoint, and that checkpoint is a feature rather than an obstacle.

Booking, paying, sending, and deleting are the obvious four. A consultation request submitted by software on somebody's behalf, without that person seeing what was sent, is a problem for the firm and for the client. The correct design lets the task get as far as a filled form and stop, with a human confirming.

Authentication and paywalls are legitimate stops too, and they should be recorded as such rather than as failures. A readiness report that lists a login wall as a blocking defect is measuring the wrong thing and will push somebody to weaken a control that exists for a reason.

Actual PageOptimized screenUse the numbered steps with the product image below.
Actual product

Keep prompt evidence and sources visible

AI Tracker stores the prompts, runs, responses, mentions, citations, competitors, and source gaps needed for a reproducible observation.

PageOptimized AI Tracker prompts view showing tracked prompts, engines, mentions, citations, competitors, and run history.Open full size
  1. Approve the prompt and protocol.
  2. Record repeated runs.
  3. Inspect the response and sources.
  4. Create only evidence-backed follow-up work.

Run a readiness check

Seven steps. The static preflight is cheap and the real task run is the evidence.

  1. Inspect what exists without interaction.The rendered information available before anything is clicked. Content that requires a click is content an agent may never reach.
  2. Check headings, landmarks, names, roles, and states.The static preflight. It is fast, it catches the majority of blocking problems, and it does not prove a task succeeds.
  3. Complete the primary path with a keyboard.If you cannot, software cannot. This single test finds more real blockers than any automated scan.
  4. Run the real task with a defined stopping point.Short of any irreversible action. The boundary is part of the test design rather than a limitation of it.
  5. Document the authentication and consent boundaries.Where a task legitimately cannot proceed. A paywall or a login is a correct stop, not a failure, and the record should say which it was.
  6. Save the ordered evidence.Steps, states, semantics, screenshots, and the failure point. This is what makes a rerun a comparison instead of a fresh opinion.
  7. File failures as experience issues.With the affected task named. They compete for priority on their real merits, which is usually enough, and they stay honest.
Worked exampleSee the completed Horizon Legal work, then build your version.
Completed example: Horizon Legal

Agentic browser task and verified rerun

The same three-step task failed before the calculator rendering fix and passed afterward, with both executions preserved.

Before this lesson: one user task and its known page failure

Task
Use the calculator and reach a consultation path
Known technical fact
Initial HTML is not readable
Saved run evidence
Steps, labels, state, errors, and result
Rerun
Failed baseline followed by a dated passing run

After this lesson: Finished output

Task
Reach the Austin settlement calculator, enter a harmless sample, and read the result without submitting personal information
Initial run
Failed at step 3: the calculator shell loaded but the result status region never appeared
Rerun
Passed after the rendering fix; the saved semantic evidence names the form, medical-cost input, and illustrative-estimate status
Comparison
The passing run links to the failed baseline and retains runner, dates, step outcomes, screenshot URL, and failure history
Decision

Call the task verified only because every required step passed in the dated rerun; do not turn that into a crawler, ranking, or citation claim.

Save this in

Site Health -> AI access and agentic browsing -> Reach and use the settlement calculator.

Your turn

Use the principle on your own project

Follow the sequence once. The goal is a defensible decision, not completing steps for their own sake.

Have these open

One primary task path · Rendered page, headings, landmarks, labels, keyboard operation, state, errors, and authentication boundaries

  1. Inspect the rendered information available without hidden interaction.
  2. Check headings, landmarks, labels, names, roles, and states.
  3. Complete the primary path with keyboard and accessibility semantics.
  4. Document authentication, paywall, consent, and destructive-action boundaries.
  5. Save every step result, observed state, semantic evidence, screenshot URL, failure, approval, and execution source in Site Health.
  6. Record failures as user-experience and automation-readiness issues, not ranking guarantees.
Reference notesDefinitions, site-specific paths, common mistakes, and completion paths

Terms in plain language

Use these definitions when a term is unfamiliar.

Agentic browsing

Software interpreting a task and interacting with a site through navigation, forms, controls, and changing states to attempt completion.

Example

Find an Austin consultation option, select an available time, enter valid details, stop before final submission, and report blockers.

Semantic control

An interface element whose role, name, state, and relationship can be understood programmatically and by assistive technology.

Example

A button named Check availability is clearer than a clickable div with no role or label.

Safe confirmation

An explicit checkpoint before an irreversible, costly, private, or externally visible action.

Example

Require confirmation before submitting a consultation, purchase, redirect, deletion, or publication.

Choose the path that matches your site

New sites establish evidence; established sites use history.

Brand-new site or no usable history

Include keyboard, labels, validation, stable states, authentication boundaries, and safe confirmation in design and QA. Run real representative tasks before launch.

Site with usable history

Use static checks to find likely barriers, then run recorded tasks through important journeys and separate site failure, agent failure, and unsupported authentication or policy boundaries.

Common mistakes

What people often do and what to do instead.

Testing only whether the page loads
InsteadComplete the task and record where meaning, controls, state, or recovery fails.
Calling an accessibility failure an AI ranking factor
InsteadReport it as user and automation readiness evidence without causal ranking claims.
Current workflow

Complete this in PageOptimized

What PageOptimized does
Site Health stores reusable browser-task definitions and immutable runs with exact definition snapshots, ordered step results, DOM or accessibility evidence, screenshot and evidence URLs, failures, approvals, execution source, reviewer, timing, and rerun baselines.
What it does not do
The saved execution source states whether the run came from an external browser agent, an integration, or a manually observed browser session. Saving evidence does not pretend PageOptimized itself drove that browser.
How to use it now
Use the checker as a static preflight, define the real task in Site Health, run it through the named browser source, save every step, and compare the post-fix rerun with the failed baseline.

You should now have

  • A step-by-step task result
  • Accessible-name, state, error, and boundary failures
  • Safe remediation and retest criteria

Before you move on, confirm

  • Important controls have accessible names.
  • Task state and errors are understandable.
  • Sensitive actions require explicit confirmation.
Lesson 10.4

Use honest evidence language

Separate an observed answer from a causal explanation.

Answer engines are non-deterministic in a way search results are not. The same prompt returns different answers depending on the engine, the model version, the product mode, the date, the location, the account, and which sources happened to be retrievable in that moment.

That does not make observation worthless. It makes the protocol the thing that carries the meaning, and it makes the language you use about the results the difference between a defensible report and one that falls apart the first time somebody checks.

One run is an observation under conditions. Say what you recorded, not what the system does.

Observation, generalization, and causal claim

An observation is what one recorded test returned under its exact conditions. A generalization is a broader claim from repeated observations, which needs enough varied evidence and still needs its boundaries stated. A causal claim says something you did produced the outcome, and timing alone never supports it.

The defensible sentence names the numbers and the scope. Horizon Legal was cited in 6 of 12 recorded Perplexity snapshots for the fixed Austin-fees prompt cohort, between these dates. Anybody can check that, and it stays true regardless of what the engine does next week.

The indefensible version is shorter and much more appealing: we made Horizon the preferred AI answer. It claims a universal outcome from a sample, attributes it to a cause nobody tested, and will be quoted back to you the first time a client runs the prompt themselves and sees something else.

Four things to classify separately

Collapsing these into a single mentioned flag is what makes AI visibility reporting useless. They move independently:

  • Mention. The brand appears in the answer.
  • Recommendation. The answer suggests it, which is not the same as naming it.
  • Citation. A source link is attached, which may not be your page and may not be where the claim came from.
  • Accuracy. What was said about you is true, which is independent of all three above.

A mention that is inaccurate is a problem rather than a win, and a citation of a third-party page that describes you wrongly is a module 7 reputation task rather than a visibility success.

Record the raw response alongside the classification. Six months on, the classification is somebody's judgment and the raw text is the evidence, and only one of those settles a disagreement.

How many runs before you can say anything

There is no published threshold, so the honest approach is to make the sample visible rather than to claim a standard. Six of twelve snapshots is a statement anybody can weigh. Cited by Perplexity is not, and it is the same underlying observation with the sample deleted.

Hold the prompt cohort fixed across runs. Comparability comes from the prompts staying identical, in the same order, under the same recorded conditions. A cohort somebody edited between runs produces a chart with a discontinuity nobody documented, and it will be read as a trend.

Expect variance and say so in advance. Running the same prompt three times in one afternoon and getting three different answers is the normal behavior of these systems, not a fault in your protocol, and a stakeholder who learns that from you before it happens responds very differently to the one who discovers it themselves.

Record an observation properly

Five fields, then a confidence and a retest date. Anything missing from the first five makes the observation uncomparable.

  1. Record the conditions before the result.Prompt text, engine, product mode, date, location, and account state. These are what make two runs comparable, and they are impossible to reconstruct afterward.
  2. Save the answer and its cited sources verbatim.The raw text, not a summary. A summary is a second interpretation layered on an unstable observation.
  3. Classify mention, recommendation, citation, and accuracy separately.Four independent facts. A single score collapses them and destroys the information you collected them for.
  4. List the plausible explanations you did not rule out.Including the possibility that the run was simply variable. This is the field that keeps the report honest and it takes one line.
  5. Assign a confidence and a retest date.A single observation is low confidence by construction. The retest is what converts it into something you can generalize from, and without a date it will not happen.
Worked exampleSee the completed Horizon Legal work, then build your version.
Completed example: Horizon Legal

Evidence-safe finding

The wording says what the test observed and stops before unsupported causation.

Before this lesson: raw observations that could be overstated

Prompt cohort
Five approved injury-law buyer questions
Engines
ChatGPT and Perplexity
Captured evidence
Responses, mentions, citations, and sources
Unknown
Why any engine selected a source

After this lesson: Finished output

Unsupported
Adding a special file will make AI engines cite us
Observed
ChatGPT cited Horizon pages on the Aug 28 run; Perplexity named Horizon on 4 of 5 prompts but cited it on none
Plausible
Owned case proof may help answer legal-evaluation prompts
Unknown
Why the engine selected it in any individual response
Decision

Report prompt-level observations and source coverage; never promise inclusion or causation.

Save this in

AI Tracker finding with prompt, engine, mode, date, response, citation, and accuracy.

Your turn

Use the principle on your own project

Follow the sequence once. The goal is a defensible decision, not completing steps for their own sake.

Have these open

Prompt, engine, mode, account state, location, date, response, and sources · Definitions for mention, recommendation, accuracy, and citation

  1. Record prompt, engine, mode, date, location, and account state.
  2. Save the answer and cited sources.
  3. Classify mention, recommendation, accuracy, and citation separately.
  4. List plausible explanations and verification steps.
  5. Assign confidence and a retest schedule.
Reference notesDefinitions, site-specific paths, common mistakes, and completion paths

Terms in plain language

Use these definitions when a term is unfamiliar.

Observation

What one recorded test actually returned under its exact engine, model, mode, prompt, date, location, and account conditions.

Example

ChatGPT named Horizon Legal second for prompt P-04 on Aug 28 in the recorded mode.

Generalization

A broader claim inferred from repeated observations. It requires enough varied evidence and still needs stated boundaries.

Example

Three stable weekly citations suggest recurring visibility for this prompt cohort, not universal preference across AI systems.

Causal explanation

A claim that a specific change produced an answer outcome. Timing or correlation alone is insufficient.

Example

Publishing a guide before a citation appeared does not prove the guide caused the citation.

Choose the path that matches your site

New sites establish evidence; established sites use history.

Brand-new site or no usable history

Record the first answer snapshots as a baseline and expect high uncertainty. Avoid success or failure claims until the prompt cohort and cadence produce comparable observations.

Site with usable history

Compare fixed prompt cohorts with source, content, technical, and release history while preserving model or product changes and normal variation.

Common mistakes

What people often do and what to do instead.

Generalizing from one answer
InsteadCall it one observation and repeat under a defined protocol.
Treating a citation as endorsement or accuracy
InsteadReview citation, claim support, accuracy, and recommendation separately.

You should now have

  • A reproducible observation
  • Separate classifications
  • Confidence, plausible explanations, and retest date

Before you move on, confirm

  • Protocol is reproducible.
  • Citation and accuracy are separate.
  • The report avoids deterministic causal claims.
Primary references

Verify the practice at the source.

Practices and source links reviewed August 2026.