ACADEMY/MODULE 06
AEO70 min guidedFieldwork: Baseline run plus monthly repetitionsBUILD: AI visibility baseline

Audit every answer engine

Measure brand inclusion, recommendation position, cited sources, accuracy, and competitor share across answer engines.

IN PLAIN ENGLISH

Run your prompt set against every engine, the same way each time, and write down what comes back, so you have a dated baseline instead of a folder of screenshots.

THE OUTCOME

You will have a reproducible baseline showing where the brand appears, which source patterns recur, and which explanations need verification.

WHY THIS MODULE MATTERS

AI answers vary by engine, model, date, location, and wording. A useful audit controls what it can, records what changed, and looks for patterns across a prompt set rather than treating one screenshot as truth.

IF YOU SKIP IT / No baseline means no provable progress. Six months from now, someone will ask what changed, and "it feels better" will not survive the budget review.

EVIDENCE STANDARD

Agreement and recurrence are observational signals

Cross-engine agreement does not prove settled consensus. A missing citation does not prove a trust gap. A recurrent source is a candidate for verification, not a proven gatekeeper.

JARGON, TRANSLATED / WORDS YOU WILL MEET IN THIS MODULE
Baseline
The recorded, dated snapshot of what every engine says before you change anything. All future proof is measured against it.
Run
One prompt asked once in one engine. The unit you record. Cohort results are built from many runs.
Citation
A source URL an engine names or links when answering. It is evidence that the URL appeared as support in that run, not proof that it caused the complete answer.
Share of voice
How often you are named across a cohort's runs versus competitors. Reported as a share, never a raw count.
Hallucination
A confident but false statement by a model, such as outdated pricing or a feature you do not have. Logged and corrected, not laughed off.
FOLLOW THIS EXACT SEQUENCE

Each phase feeds the next. Don't skip ahead, the artifact at the end is only trustworthy if every phase ran.

  1. 01Control

    Fix prompt wording, engine, surface, account state, geography, and run cadence.

  2. 02Capture

    Store the complete answer, cited URLs, date, and screenshots where policy permits.

  3. 03Label

    Mark absent, mentioned, cited, compared, recommended, or incorrectly represented.

  4. 04Score

    Measure accuracy, prominence, sentiment, citation quality, and competitive position separately.

  5. 05Explain

    List plausible explanations, supporting and conflicting evidence, verification steps, and confidence.

  6. 06Assign

    Choose an owned-content, source, entity, product, or measurement action with an owner.

  7. 07Retest

    Use a fixed panel and rolling panel so you can separate trend from model volatility.

  8. You now have

    AI visibility baseline

PART 1 LEARN THE METHOD
06.1

Define the test protocol

Consistency makes snapshots comparable.

Scientists keep lab notebooks for a reason: an experiment you cannot repeat is an anecdote. AI answers drift between engines, days, accounts, and even runs of the same prompt, so casual testing produces nothing but anecdotes with screenshots.

The protocol is your notebook. Same prompts, same engines, fresh threads, three to five repetitions, recorded the same way, and record everything: full answer text, every cited URL, the model label, the date. The citations you save today become Module 08's target list, and there is no going back for them later. Yes, this is tedious. Tedious is what reproducible feels like on day one.

HELD CONSTANT
  • ENGINESThe same named list every run
  • ACCOUNT STATETemporary chat, memory off, logged out
  • LOCATIONNoted and unchanged
  • PROMPT ORDERSame order, fresh thread each time
  • REPETITIONS3 to 5 runs of every prompt
  • CAPTUREThe same columns, every time
THE ONLY THING THAT MOVES
The prompt

One dimension at a time, taken from the frozen set you built in Module 05.

DRIFTS ON ITS OWN

Answers vary between engines, days, accounts, and repeat runs of the identical prompt. That is why repetitions are part of the protocol, and why one run proves nothing.

EVERY RUN LANDS IN THE SAME ROW
RUN IDDATEENGINEPROMPTFULL ANSWERCITATION URLSSCREENSHOT

The citation URLs you save today become the Module 08 target list. There is no going back for them later.

WHAT THIS SHOWS / An experiment you cannot repeat is an anecdote with a screenshot attached. The protocol fits on one page, and the real test is whether someone else could run it without you.
DO THIS, IN ORDER / 3 STEPS
  • 01

    Choose engines, account state, location, date, prompt set, and number of repetitions.

  • 02

    Use fresh threads and run prompts in the same order when possible.

  • 03

    Record the complete answer, cited URLs, and model label; do not reduce results to a yes/no mention.

WHERE TO DO THIS, EXACTLY
  • Write the protocol on one page: engines (ChatGPT, Perplexity, Gemini, Claude, Copilot, AI Overviews), logged-out or temporary chats, memory off, location noted, 3-5 repetitions, same prompt order.
  • Create the capture sheet with columns: run ID, date, engine, prompt, full answer (pasted), citation URLs, screenshot link. Store screenshots in one dated folder.
  • Block half a day and run the frozen set. Tedious is correct; tedious is what makes it reproducible.
CHECKPOINT / HOW YOU KNOW IT WORKED

The protocol fits on one page, someone else could run it without you, and every run produced a saved answer with saved citations.

  • NO BUDGET?

    Manual is fine to start. When the set grows past ~40 prompts, a tracker automates the capture and the protocol stays the same either way. On WordPress, CiteTrack AI runs it from the dashboard you already have. Everywhere else, Pageoptimized does the same job as a platform. Both are ours, and there are other tools worth comparing.

  • WATCH OUT

    A failed run and a genuine absence look identical once exported. Record the response state beside every result: valid answer, brand absent, empty response, retry, terminal failure. Counting a failed run as a zero is how a measurement program starts reporting fiction.

  • WATCH OUT

    A logged-in account with chat memory contaminates every answer. Temporary chats or a clean profile, always.

  • PRO TIP

    In ChatGPT, note whether the answer used web search or model memory. The two produce different citation behavior and are worth tracking separately.

OPTIONAL / AI CO-PILOTRun this lesson with AI, the prompt, sources and expected output
HUMAN OWNSJudgment and approval

Choose the business priority, interpret exceptions, protect confidential data, challenge weak evidence, and approve the final decision.

AI ASSISTSAnalysis and structure

Clean, classify, compare, calculate, and draft rows from supplied evidence. AI may surface patterns; it does not own strategy.

SOURCE FILESVisibility dataset

Citation tracker, plus answer captures from ChatGPT / Gemini / Perplexity.

OUTPUT PREVIEW / answer-capture-log.CSVOne example row from the finished deliverable
EXAMPLE DATA - REPLACE IT
PromptEngine / surfaceDateFull answer savedCitations savedLocationRun ID
Best agency reporting softwareChatGPT Search2026-08-07Yes4USAEO-0241
CLAUDE / CHATGPT PROMPTAnalyze the evidence without outsourcing the decision
BEFORE RUNNING

Remove personal or confidential data. Attach the named exports, explain every column, and tell the model when the dataset was collected.

You are assisting a human B2B SaaS SEO and AEO operator with Module 06: Audit every answer engine.

Current lesson: Define the test protocol
Objective: Consistency makes snapshots comparable.
Required artifact: AI visibility baseline

BUSINESS CONTEXT I WILL PROVIDE
- Product, category, target market, pricing model, sales motion, and geography
- The priority customer segment and the commercial outcome for this 90-day cycle
- Visibility dataset exported from Citation tracker
- Answer captures exported from ChatGPT / Gemini / Perplexity
- Definitions for any internal fields, stages, scores, and abbreviations

TASK
1. Choose engines, account state, location, date, prompt set, and number of repetitions.
2. Use fresh threads and run prompts in the same order when possible.
3. Record the complete answer, cited URLs, and model label; do not reduce results to a yes/no mention.

REQUIRED OUTPUT
Return a table using these exact columns: Prompt | Engine / surface | Date | Full answer saved | Citations saved | Location | Run ID.
For every recommendation, cite the source row, URL, call note, or data point that supports it.
Add a confidence column in your analysis: High, Medium, or Low.
List missing evidence separately instead of guessing.
Finish with a section named HUMAN DECISIONS REQUIRED.

RULES
- Do not invent search volume, revenue, customer statements, product capabilities, or competitor facts.
- Do not treat correlation as causation.
- Preserve contradictory evidence and explain why it conflicts.
- Do not make the final priority or publishing decision. Prepare the evidence for a human owner.
- Use this example only as a format reference, not as evidence: Run 40 prompts across ChatGPT, Perplexity, Gemini, and Google AI experiences on the same two-day window.
HUMAN APPROVAL GATE

The test protocol is documented. A named human owner must verify this before the lesson is complete.

B2B SAAS EXAMPLE

Run 40 prompts across ChatGPT, Perplexity, Gemini, and Google AI experiences on the same two-day window.

06.2

Score visibility and accuracy

A mention can still be unhelpful or wrong.

On a live engagement we watched Claude name the brand on 76% of prompts while ChatGPT managed 26%. Same company, same week, three-fold spread. Average those into one number and you have hidden the most useful finding in the audit.

So score dimensions separately: present or not, position, recommended or merely mentioned, sentiment, accuracy. The gap between appearing and being recommended deserves special attention. Showing up in 35% of answers but recommended in 12% is a precise diagnosis: the engines know you exist and cannot find a reason to pick you. That reason is called comparative proof, and Modules 03 and 07 build it.

ONE BRAND, ONE WEEK, TWO ENGINES
CLAUDE76%
CHATGPT26%
BLENDED51%

A threefold spread, reported as one number nobody experienced. Averaging the engines deletes the most useful finding in the audit.

PRESENT IS NOT RECOMMENDED
NAMED35%
RECOMMENDED12%
23 POINTS

The engines know you exist and are not choosing you. That is an observed signal worth investigating, not a diagnosis.

BEFORE YOU ACT ON THE GAP, CHECK
COMPARATIVE PROOFPRODUCT FITPROMPT INTERPRETATIONSOURCE COVERAGERUN VARIANCE
WHAT THIS SHOWS / Presence, position, recommendation, sentiment, and accuracy are separate columns because each one is a different problem with a different owner. A blended score hides all of them.
DO THIS, IN ORDER / 3 STEPS
  • 01

    Score presence, recommendation position, sentiment, feature accuracy, and citation ownership.

  • 02

    Mark hallucinated claims, outdated pricing, wrong positioning, and missing best-fit context.

  • 03

    Calculate share of answers and share of cited sources by prompt tier and engine.

WHERE TO DO THIS, EXACTLY
  • Add label columns per run: brand present (yes/no), position in list, recommended (yes/no), sentiment (positive/neutral/negative), accuracy (accurate/partial/wrong).
  • Log each factual error as its own row in an accuracy tab: the claim, the truth, the likely stale source.
  • Compute rates with a pivot table: presence and recommendation rates per cohort per engine. Twenty minutes of spreadsheet work, no code.
CHECKPOINT / HOW YOU KNOW IT WORKED

Every run is scored for presence, position, sentiment, and accuracy separately, and every factual error is logged as its own line item.

  • PRO TIP

    The presence-versus-recommendation gap is a useful observed signal, not a diagnosis. Investigate comparative proof, product fit, prompt interpretation, source coverage, and run variance before choosing an action.

  • WATCH OUT

    Do not average sentiment across engines into one number. Claude being warm and Gemini being cold is a finding, not noise to smooth.

OPTIONAL / AI CO-PILOTRun this lesson with AI, the prompt, sources and expected output
HUMAN OWNSJudgment and approval

Choose the business priority, interpret exceptions, protect confidential data, challenge weak evidence, and approve the final decision.

AI ASSISTSAnalysis and structure

Clean, classify, compare, calculate, and draft rows from supplied evidence. AI may surface patterns; it does not own strategy.

SOURCE FILESAnswer captures

ChatGPT / Gemini / Perplexity, plus source authority context from Ahrefs / Semrush.

OUTPUT PREVIEW / brand-state-labels.CSVOne example row from the finished deliverable
EXAMPLE DATA - REPLACE IT
Run IDBrand statePositionSentimentRecommendationAccuracyReviewer
AEO-0241Compared2PositiveYesPartialAEO analyst
CLAUDE / CHATGPT PROMPTAnalyze the evidence without outsourcing the decision
BEFORE RUNNING

Remove personal or confidential data. Attach the named exports, explain every column, and tell the model when the dataset was collected.

You are assisting a human B2B SaaS SEO and AEO operator with Module 06: Audit every answer engine.

Current lesson: Score visibility and accuracy
Objective: A mention can still be unhelpful or wrong.
Required artifact: AI visibility baseline

BUSINESS CONTEXT I WILL PROVIDE
- Product, category, target market, pricing model, sales motion, and geography
- The priority customer segment and the commercial outcome for this 90-day cycle
- Answer captures exported from ChatGPT / Gemini / Perplexity
- Source authority context exported from Ahrefs / Semrush
- Definitions for any internal fields, stages, scores, and abbreviations

TASK
1. Score presence, recommendation position, sentiment, feature accuracy, and citation ownership.
2. Mark hallucinated claims, outdated pricing, wrong positioning, and missing best-fit context.
3. Calculate share of answers and share of cited sources by prompt tier and engine.

REQUIRED OUTPUT
Return a table using these exact columns: Run ID | Brand state | Position | Sentiment | Recommendation | Accuracy | Reviewer.
For every recommendation, cite the source row, URL, call note, or data point that supports it.
Add a confidence column in your analysis: High, Medium, or Low.
List missing evidence separately instead of guessing.
Finish with a section named HUMAN DECISIONS REQUIRED.

RULES
- Do not invent search volume, revenue, customer statements, product capabilities, or competitor facts.
- Do not treat correlation as causation.
- Preserve contradictory evidence and explain why it conflicts.
- Do not make the final priority or publishing decision. Prepare the evidence for a human owner.
- Use this example only as a format reference, not as evidence: The brand appears in 35% of Tier 1 answers but is recommended in only 12%. Treat weak comparative proof, prompt fit, product fit, source coverage, and sampling variance as hypotheses to verify.
HUMAN APPROVAL GATE

Answers and citations are stored, not just scores. A named human owner must verify this before the lesson is complete.

B2B SAAS EXAMPLE

The brand appears in 35% of Tier 1 answers but is recommended in only 12%. Treat weak comparative proof, prompt fit, product fit, source coverage, and sampling variance as hypotheses to verify.

06.3

Investigate cited sources

Source recurrence is an observed signal, not proof of influence.

When an engine answers, ask whose homework it was copying. The cited sources are the answer, and they repeat. Classify every cited URL by type, count recurrences, and a short list of gatekeepers appears: the same category roundup, the same review page, the same comparison, cited across engines again and again.

Read those pages the way a model reads them. What do they claim about you, how fresh are they, are you even in them? One stale roundup cited by three engines can matter more than your next ten blog posts, which is a strange thing to accept and a very cheap thing to fix.

DO THIS, IN ORDER / 3 STEPS
  • 01

    Group cited URLs by first-party, review, media, community, video, competitor, and unknown.

  • 02

    Count recurrence across engines, prompt archetypes, and repeated runs.

  • 03

    Inspect recurrent pages for claims, entities, freshness, relevance, and accuracy; retest before assigning causality.

WHERE TO DO THIS, EXACTLY
  • Copy every citation URL from the capture sheet into a sources tab, one row per citation, and classify the domain: yours, review platform, listicle, media, community, competitor, other.
  • Use COUNTIF to rank domains and URLs by recurrence across the sampled runs. Label the top repeaters as candidates for verification, then add buyer relevance, run coverage, freshness, and confidence before choosing an action.
  • Open the top five cited pages and read them as a model would: what claims they make about you and competitors, when they were last updated, whether you are present and described correctly.
CHECKPOINT / HOW YOU KNOW IT WORKED

Cited sources are classified and counted, and you can name the five URLs that most influence your category's answers.

  • PRO TIP

    Record the exact sentence each source says about you. In Module 08 you will pitch corrections, and you cannot correct what you did not write down.

  • WATCH OUT

    Engines sometimes cite URLs that do not exist or now 404. Verify each top source loads before you build strategy on it.

OPTIONAL / AI CO-PILOTRun this lesson with AI, the prompt, sources and expected output
HUMAN OWNSJudgment and approval

Choose the business priority, interpret exceptions, protect confidential data, challenge weak evidence, and approve the final decision.

AI ASSISTSAnalysis and structure

Clean, classify, compare, calculate, and draft rows from supplied evidence. AI may surface patterns; it does not own strategy.

SOURCE FILESSource authority context

Ahrefs / Semrush, plus impact overlay from GSC / analytics.

OUTPUT PREVIEW / citation-source-map.CSVOne example row from the finished deliverable
EXAMPLE DATA - REPLACE IT
Prompt familyCited URLDomain classClaims supportedBrand presentCompetitor presentOpportunity
Agency reportingg2.com/categories/...ReviewCategory shortlistYesYesImprove review proof
CLAUDE / CHATGPT PROMPTAnalyze the evidence without outsourcing the decision
BEFORE RUNNING

Remove personal or confidential data. Attach the named exports, explain every column, and tell the model when the dataset was collected.

You are assisting a human B2B SaaS SEO and AEO operator with Module 06: Audit every answer engine.

Current lesson: Investigate cited sources
Objective: Source recurrence is an observed signal, not proof of influence.
Required artifact: AI visibility baseline

BUSINESS CONTEXT I WILL PROVIDE
- Product, category, target market, pricing model, sales motion, and geography
- The priority customer segment and the commercial outcome for this 90-day cycle
- Source authority context exported from Ahrefs / Semrush
- Impact overlay exported from GSC / analytics
- Definitions for any internal fields, stages, scores, and abbreviations

TASK
1. Group cited URLs by first-party, review, media, community, video, competitor, and unknown.
2. Count recurrence across engines, prompt archetypes, and repeated runs.
3. Inspect recurrent pages for claims, entities, freshness, relevance, and accuracy; retest before assigning causality.

REQUIRED OUTPUT
Return a table using these exact columns: Prompt family | Cited URL | Domain class | Claims supported | Brand present | Competitor present | Opportunity.
For every recommendation, cite the source row, URL, call note, or data point that supports it.
Add a confidence column in your analysis: High, Medium, or Low.
List missing evidence separately instead of guessing.
Finish with a section named HUMAN DECISIONS REQUIRED.

RULES
- Do not invent search volume, revenue, customer statements, product capabilities, or competitor facts.
- Do not treat correlation as causation.
- Preserve contradictory evidence and explain why it conflicts.
- Do not make the final priority or publishing decision. Prepare the evidence for a human owner.
- Use this example only as a format reference, not as evidence: Three engines repeatedly cite the same category roundup. Inspect it and buyer usage, correct inaccuracies where possible, then retest instead of assuming it controls the shortlist.
HUMAN APPROVAL GATE

Accuracy problems are separated from visibility gaps. A named human owner must verify this before the lesson is complete.

B2B SAAS EXAMPLE

Three engines repeatedly cite the same category roundup. Inspect it and buyer usage, correct inaccuracies where possible, then retest instead of assuming it controls the shortlist.

06.4

Turn gaps into actions

The audit is valuable only when it changes the roadmap.

An audit that ends in a slide deck was a month of work donated to a filing cabinet. The version that pays rent converts every gap into a classified, owned action: content gaps become pages, accuracy gaps become corrections, source gaps become outreach, each with one owner and one retest date.

Then re-run monthly, same protocol, and judge cohorts instead of single prompts. Win 7 of 10 in a cohort and you have won the cohort. That sentence, backed by a dated baseline, is what survives a budget review. A folder of screenshots is not.

DO THIS, IN ORDER / 3 STEPS
  • 01

    Classify each gap as content, entity, evidence, source, product-fit, or measurement.

  • 02

    Assign one owner and one verification date.

  • 03

    Re-run the stable prompt set monthly and research prompts after major changes.

WHERE TO DO THIS, EXACTLY
  • Move every gap into an action queue tab: finding, root cause, action type (content, correction, source), owner, retest date.
  • Put the monthly re-run in the calendar now, same protocol, same set, and add its date to the sheet.
  • Report cohort movement, not prompt movement, when you share results internally.
CHECKPOINT / HOW YOU KNOW IT WORKED

Every priority gap has an owner and a retest date, and next month's re-run is already scheduled with the same protocol.

  • PRO TIP

    Accuracy corrections are the fastest wins in the whole program: fix the stale source, and the wrong claim often disappears within weeks.

  • WATCH OUT

    Do not chase every gap. Tier them by the Module 05 scores and work the top of the list only.

OPTIONAL / AI CO-PILOTRun this lesson with AI, the prompt, sources and expected output
HUMAN OWNSJudgment and approval

Choose the business priority, interpret exceptions, protect confidential data, challenge weak evidence, and approve the final decision.

AI ASSISTSAnalysis and structure

Clean, classify, compare, calculate, and draft rows from supplied evidence. AI may surface patterns; it does not own strategy.

SOURCE FILESImpact overlay

GSC / analytics, plus visibility dataset from Citation tracker.

OUTPUT PREVIEW / aeo-action-queue.CSVOne example row from the finished deliverable
EXAMPLE DATA - REPLACE IT
FindingRoot causeAction typeAsset / sourceOwnerRetest dateSuccess state
Absent from workflow promptsNo direct guideOwned content/automated-reporting/Content2026-09-01Mentioned or cited
CLAUDE / CHATGPT PROMPTAnalyze the evidence without outsourcing the decision
BEFORE RUNNING

Remove personal or confidential data. Attach the named exports, explain every column, and tell the model when the dataset was collected.

You are assisting a human B2B SaaS SEO and AEO operator with Module 06: Audit every answer engine.

Current lesson: Turn gaps into actions
Objective: The audit is valuable only when it changes the roadmap.
Required artifact: AI visibility baseline

BUSINESS CONTEXT I WILL PROVIDE
- Product, category, target market, pricing model, sales motion, and geography
- The priority customer segment and the commercial outcome for this 90-day cycle
- Impact overlay exported from GSC / analytics
- Visibility dataset exported from Citation tracker
- Definitions for any internal fields, stages, scores, and abbreviations

TASK
1. Classify each gap as content, entity, evidence, source, product-fit, or measurement.
2. Assign one owner and one verification date.
3. Re-run the stable prompt set monthly and research prompts after major changes.

REQUIRED OUTPUT
Return a table using these exact columns: Finding | Root cause | Action type | Asset / source | Owner | Retest date | Success state.
For every recommendation, cite the source row, URL, call note, or data point that supports it.
Add a confidence column in your analysis: High, Medium, or Low.
List missing evidence separately instead of guessing.
Finish with a section named HUMAN DECISIONS REQUIRED.

RULES
- Do not invent search volume, revenue, customer statements, product capabilities, or competitor facts.
- Do not treat correlation as causation.
- Preserve contradictory evidence and explain why it conflicts.
- Do not make the final priority or publishing decision. Prepare the evidence for a human owner.
- Use this example only as a format reference, not as evidence: Incorrect integration claim -> update product docs, schema/entity references, partner profile, and the cited third-party listing.
HUMAN APPROVAL GATE

Every priority gap has an owner and retest date. A named human owner must verify this before the lesson is complete.

B2B SAAS EXAMPLE

Incorrect integration claim -> update product docs, schema/entity references, partner profile, and the cited third-party listing.

06.5

Interpret patterns without inventing causes

The audit produces observations, not automatic diagnoses. Agreement, absence, and recurring citations narrow the investigation but do not prove why an answer appeared.

A hung jury can still be persuaded; a unanimous one has moved on. Engine agreement works the same way. When every engine names the same competitor for an attribute, the web has reached consensus and moving it is slow, compounding work. When engines split 2-2, nothing is settled, and a few well-placed sources can tip the verdict. The splits are your attack list.

And file all three outputs of the audit, because most teams only file one. Opportunities get the attention. Objections, the recurring negatives AI repeats about you, get ignored until a deal dies of one. Strengths get taken for granted until the source keeping them alive goes stale and they quietly vanish. Objections get their playbook in Module 08. Strengths get protected, which is cheaper than winning them back.

DO THIS, IN ORDER / 3 STEPS
  • 01

    Write each finding as: observed signal, plausible explanations, verification needed, confidence, and reversible next action.

  • 02

    Tag recurring negative or hedged statements by theme, then verify each against product facts, cited sources, prompt wording, and repeated runs before treating it as a market objection.

  • 03

    Record accurate favorable patterns and their cited sources as monitoring candidates; do not claim one source caused or maintains the result without a controlled retest.

WHERE TO DO THIS, EXACTLY
  • Add agreement columns per attribute: engines tested, repetitions, engines naming each brand, run date, and account/location conditions. Describe the result as high, mixed, or low agreement in this sample.
  • Create an objections tab: every negative or hedged phrase, tagged by theme, with a count. Three or more mentions makes the objection list.
  • Create a strengths tab: attributes where sampled answers are accurate and favorable, with observed citations and a verification plan. Do not assign causality to one source without a retest.
CHECKPOINT / HOW YOU KNOW IT WORKED

Each Tier 1 attribute records the engine split, repetitions, plausible explanations, verification status, and confidence; objections and strengths remain observed patterns until validated.

  • PRO TIP

    Mixed results plus high commercial value can justify a small experiment, but prioritize only after checking product fit, source overlap, effort, and confidence.

  • WATCH OUT

    More engines and repeated runs improve confidence, but no fixed engine count turns observational agreement into causal proof.

OPTIONAL / AI CO-PILOTRun this lesson with AI, the prompt, sources and expected output
HUMAN OWNSJudgment and approval

Choose the business priority, interpret exceptions, protect confidential data, challenge weak evidence, and approve the final decision.

AI ASSISTSAnalysis and structure

Clean, classify, compare, calculate, and draft rows from supplied evidence. AI may surface patterns; it does not own strategy.

SOURCE FILESVisibility dataset

Citation tracker, plus answer captures from ChatGPT / Gemini / Perplexity.

OUTPUT PREVIEW / audit-outputs.CSVOne example row from the finished deliverable
EXAMPLE DATA - REPLACE IT
AttributeEngine agreementStateOutput typeThemeRecurrenceRoute
Agency reporting2 of 4 name brandUnsettledOpportunityn/an/aModule 08 sources
CLAUDE / CHATGPT PROMPTAnalyze the evidence without outsourcing the decision
BEFORE RUNNING

Remove personal or confidential data. Attach the named exports, explain every column, and tell the model when the dataset was collected.

You are assisting a human B2B SaaS SEO and AEO operator with Module 06: Audit every answer engine.

Current lesson: Interpret patterns without inventing causes
Objective: The audit produces observations, not automatic diagnoses. Agreement, absence, and recurring citations narrow the investigation but do not prove why an answer appeared.
Required artifact: AI visibility baseline

BUSINESS CONTEXT I WILL PROVIDE
- Product, category, target market, pricing model, sales motion, and geography
- The priority customer segment and the commercial outcome for this 90-day cycle
- Visibility dataset exported from Citation tracker
- Answer captures exported from ChatGPT / Gemini / Perplexity
- Definitions for any internal fields, stages, scores, and abbreviations

TASK
1. Write each finding as: observed signal, plausible explanations, verification needed, confidence, and reversible next action.
2. Tag recurring negative or hedged statements by theme, then verify each against product facts, cited sources, prompt wording, and repeated runs before treating it as a market objection.
3. Record accurate favorable patterns and their cited sources as monitoring candidates; do not claim one source caused or maintains the result without a controlled retest.

REQUIRED OUTPUT
Return a table using these exact columns: Attribute | Engine agreement | State | Output type | Theme | Recurrence | Route.
For every recommendation, cite the source row, URL, call note, or data point that supports it.
Add a confidence column in your analysis: High, Medium, or Low.
List missing evidence separately instead of guessing.
Finish with a section named HUMAN DECISIONS REQUIRED.

RULES
- Do not invent search volume, revenue, customer statements, product capabilities, or competitor facts.
- Do not treat correlation as causation.
- Preserve contradictory evidence and explain why it conflicts.
- Do not make the final priority or publishing decision. Prepare the evidence for a human owner.
- Use this example only as a format reference, not as evidence: Four engines naming the same competitor is high agreement in this sample. Verify source overlap, query interpretation, product fit, and repeatability before deciding whether to deprioritize or intervene.
HUMAN APPROVAL GATE

Every priority gap has an owner and retest date. A named human owner must verify this before the lesson is complete.

B2B SAAS EXAMPLE

Four engines naming the same competitor is high agreement in this sample. Verify source overlap, query interpretation, product fit, and repeatability before deciding whether to deprioritize or intervene.

WHAT YOU'RE LOOKING AT

The screen below shows one audited answer, annotated: the brand state, the position, the cited sources, and the accuracy verdict. This is the unit of work in this module. You will produce one of these per prompt, per engine, per run. Boring on purpose. Baselines are supposed to be boring.

YOU

Recommend reporting platforms for a 15-person marketing agency that needs white-label dashboards and HubSpot data.

ASSISTANT

For a 15-person agency on HubSpot, three platforms come up most consistently:

  1. Competitor A, strongest white-label options and agency templates.
  2. Competitor B, deep HubSpot sync, from $49/mo.
  3. Competitor C, better for larger teams.

Your brand appears in 1 of 3 runs of this prompt, never above position 3.

SOURCEScategory-roundup.comg2.comreddit.com/r/agenciescompetitor-a.com/comparison
WHAT TO LOOK FOR / Score four things on every run: are you present, at what position, is the claim about you accurate, and which sources built the answer. Store the whole answer, not a yes/no.
Real AI share-of-voice dashboard from Pageoptimized: which brands the four engines name for Swydo's category, with per-engine mention and cite rates
REAL RUN, NOT A SIMULATION / This is the same audit from Pageoptimized, our own platform, on a live engagement (July 2026): per-engine mention and cite rates for Swydo's category. Note the spread, Claude names the brand on 76% of prompts, ChatGPT on 26%. That spread is why the protocol samples every engine.
PART 2 PRACTICE IN THE EXAMPLE WORKSPACE
WHAT YOU'RE LOOKING AT

The workspace below is a filled-in audit: the capture log with run IDs, the brand-state labels, and the citation source map. The detail that matters most is the last column of the source map. Every cited URL has been turned into an opportunity or an action, which is the whole reason the audit exists.

AI visibility workspace

Answer engine audit

Measure whether the brand appears, how accurately it is described, which competitors win, and what sources shape the answer.

LOADING SAVED WORKDemo data: 24 priority prompts, weekly capture
Mention rate46%+8 pts / 30d
Citation rate21%5 owned citations
Recommendation rate29%7 shortlist wins
Answer accuracy74%6 claims to fix
VISUAL ANALYSIS

Visibility by answer surface

Separate mentions, citations, and recommendations; they are not interchangeable.

ChatGPT+12 pts
63
Google AI+5 pts
41
Perplexity+9 pts
52
Gemini-2 pts
34
Copilot+3 pts
27
0visibility rate63
PART 3 PROVE IT'S DONE
DEFINITION OF DONE0% complete
0/4

Check each item only after the artifact meets the standard. Progress saves on this device; use the academy-home backup to move it elsewhere.

YOUR NEXT STEP

You now have a dated, reproducible AI visibility baseline. Module 07 improves the clarity and evidence of priority page sections, then measures whether search and answer outcomes change.