When ChatGPT or Gemini answers a buying question with web search on, it does not search for your category. It searches for what it already remembers about your category. The queries it fans out are written from memory, before a single page is retrieved.
We ran into this in client audits before we had a name for it. A B2B SaaS with quotable pages, present on the right review sites, technically clean, still absent from the answer. Not outranked. Never searched for. It lost two stages before retrieval started, and every fix we could have sold them would have applied to the wrong stage.
Our own 101-question benchmark across four engines measured which sources get cited, which is stage four of the four below. It could not explain the brands that never entered the running at all. Reading the fan-out queries did, and once we started running the memory test in this article on client categories, the pattern was consistent enough to plan around. Three independent lines of evidence now back it: Google documents the fan-out stage itself, public benchmarking sizes the effect at roughly 3×, and controlled experiments reproduce it with the brand name as the only variable.
What follows is the operator's version of it: the mechanism behind the numbers, how memory actually gets made and on what clock, an hour-long test that tells you whether your category is memory-bound, and the two-track plan that follows from the answer. Most coverage of this effect stops at “build category authority.” That is where this starts.
How a grounded AI answer is actually assembled
Four stages. Most AEO advice starts at the third.
- Recall. For a category it has seen enough of in training, the model can name a ranked set of brands without looking anything up.
- Fan-out. It writes the search queries it intends to run, and that recall shapes them.
- Retrieval. A search backend returns pages for those queries.
- Answer. The model writes from what came back, weighted by what it already believed.
Stage four is where citation tracking lives, and it is the only stage most teams measure. Stage two is where the shortlist is really set. Put them together and the puzzle resolves: you can be perfectly retrievable and never be retrieved, because the query that would have found you was never written.
What the data shows, and what it does not
The numbers in this piece come from the largest public run of this experiment we know of: 3,960 grounded answers across 66 US buyer questions and nine industries, five runs a day for twelve days, with 13,281 fan-out queries recorded. Memory was measured by asking a model to rank the top ten brands in a category before it touched the web. Search was measured by logging every query the model then fanned out. A brand counted as searched if it appeared in any of them.
Worth saying plainly: none of that is proprietary. The protocol is two prompts and a spreadsheet, which is why we can reproduce the core of it on any client category in an afternoon and why the test later in this article is the same experiment at small scale. Cite the big run for the magnitudes. Run your own for the only numbers that decide your budget, which are yours.
| Finding | The number | What it means for you |
|---|---|---|
| Remembered brands get searched far more | 55.7% search rate when in the model's top 10, 17.4% when not | Retrieval optimization has a ceiling set upstream of retrieval. |
| Rank inside memory matters | Top 5: 67%. Ranks 6 to 10: 39%. Not in memory: 17% | Being “known” is not the goal. Being in the first five the model reaches for is. |
| Most fan-out is generic | 69% category queries, 31% brand-named; 63% of the named ones are top-5 brands | The generic 69% is your way in if you are not remembered. They are ordinary search queries, and pages can rank for them. |
| It holds in every industry tested | Not-remembered rates of 9 to 23%; remembered rates of 41 to 82% across nine verticals, including business software | This is not a consumer-brand quirk. B2B categories show the same shape. |
Now the honest part, which that run states itself and which most people re-sharing the headline will skip. It is an association in exploratory data, not a proven cause. Brand prominence is the obvious confound: famous brands are both more likely to be remembered and more likely to be searched, so some of the effect is fame, not memory feeding search. It is two models over twelve days, US prompts only, and some industries rest on six to twelve prompts. Memory was measured on a different model than search, which makes the result harder to dismiss as one model's habit but also means the exact mechanism is not pinned down. The direction is robust. The magnitudes are a floor, not a law.
A single run is a single run, so we went looking for what else supports it and what argues with it.
What corroborates this, and what argues with it
Google documents the mechanism itself. Stage two is not an inference from watching a black box. When AI Mode launched, Google described it as using “a ‘query fan-out’ technique, issuing multiple related searches concurrently across subtopics and multiple data sources”. The model writes its own queries, and the vendor says so in plain language. That settles that the stage exists. The open question is only what shapes the queries.
Experimental work isolates the brand effect that observational data cannot. The strongest objection above is that fame drives both recall and searching, so you cannot separate them by watching. A June 2026 paper by Chu and Hou attacks it the other way, by holding product specifications constant and swapping only the brand name. Across 670 valid trials where the options were identical on paper, the real brand was recommended 100% of the time. That is an experiment, not a correlation, and it isolates precisely the variable the big observational run admits it cannot.
Two more results from that paper matter more for planning than the headline does. When genuine quality information was available, product attributes explained 82.4% of the variance in what got recommended and brand identity only 1.2%. And under retrieval-augmented generation, the incumbent advantage largely collapsed. Read those together and the two-track plan later in this piece falls out of the evidence rather than out of our opinion: memory decides when the model has nothing else to go on, and evidence in the retrieved context overrides it.
The size of the override is the practical gift. In that study the brand advantage broke with a 0.075-star rating edge, 1.6 times the review count, or a 7.3% lower price. For a B2B SaaS that is not an abstraction. It means a visible, current, specific G2 profile is not a nice-to-have sitting next to your AEO work. It is the cheapest known way to beat a better-remembered competitor at the moment of comparison.
The mentions signal points the same direction. Ahrefs' 75,000-brand study found brand web mentions correlate with AI Overview visibility at 0.664 versus 0.218 for backlinks. If recall is built from corroborated mentions, that is the fingerprint you would expect to find, and it is measured on a different surface by a different party.
And the honest counterweight. The literature is not unanimous. Lichtenberg, Buchholz and Schwöbel found that an LLM recommender showed less popularity bias than conventional recommender systems on movie recommendations, with no mitigation applied. Different domain, no retrieval step, consumer catalogue rather than B2B category. But it is a real result pointing the other way, and anyone telling you models simply favour the famous in all conditions has not read it.
What survives all of it is narrower than the headline and more useful. The fan-out stage is documented. A known-brand advantage at that stage is measurable observationally and reproducible experimentally. Retrieved evidence beats it. That is enough to plan against, and it is the whole argument for running both tracks instead of picking one.
The part nobody explains, how memory gets made
A model's memory of your category is a statistical residue of its training data: how often your brand appeared, how consistently it was described, and how tightly it co-occurred with the category term and with the other brands the model already associates with that category. Three properties of that process matter more than any tactic.
- It is built from corroboration, not from your site. One page saying you are a “revenue intelligence platform” is a claim. Fifty independent pages saying it, in reviews, roundups, community threads, docs and press, is a fact the model can hold. This is the same signal behind Ahrefs' finding that brand mentions correlate with AI visibility at 0.664 versus 0.218 for backlinks. Mentions are training data. Links mostly are not.
- It is built from co-occurrence. Watch a ride-hailing fan-out and you see Waymo searched alongside Uber and Lyft, not because anyone asked about it but because the model files it on the same shelf. A brand that keeps appearing next to the top five in comparison content inherits some of their category association. Proximity is a signal.
- It updates on the training clock, not yours. Memory is frozen at a model's training cutoff and refreshes when the next generation ships, which in 2026 means a lag of months between earning a mention and the model recalling it. Everything you do for memory today is an investment in the model your buyers will be using next year. Search-led wins are the bridge until then.
One inference is ours, drawn from the client categories we have tested rather than from anyone's published run. Business software categories are narrow, and narrow categories have thinner memory: fewer brands the model recalls with confidence, and a lower bar to get into the top five. The public data points the same way, since its clearest example of an unremembered brand still getting searched came from a specialized software category. Read that two ways at once. B2B SaaS is more search-led than consumer categories, so live content can still win this quarter, and the memory slots are cheaper to take now than they will ever be again.
Which regime is your category in? Test it in an hour
There are two regimes. In a memory-bound category, the brand-named queries only ever name remembered brands and an unremembered brand almost never surfaces. In a search-led category, the generic queries still pull in brands the model did not recall. You need to know which one you are in before you spend a dollar, and you can find out today with nothing but the chat interfaces.
- Open a temporary chat in ChatGPT (the memory feature off, search off), a fresh Gemini chat, and Claude. Use a browser profile that has never discussed your brand; the personalization features will otherwise flatter you.
- Paste this, filling in your category and segment: “Without searching the web, list the 10 [category] tools you would consider for a [segment] company. Rank them by how confident you are, most confident first. Answer only from what you already know.”
- Run it 10 times per model in fresh chats. In a sheet, record for each run whether your brand appeared and at what rank. Do the same for your two closest competitors while you are there.
- Compute two numbers per model: appearance rate (runs where you were named, out of 10) and mean rank when named. Appearance rate under 30% means you are effectively not in memory. Mean rank 1 to 5 with high appearance is the top tier that gets searched about 67% of the time.
- Same question, search on. In Gemini, expand the “Searched for” or sources panel under the answer; in ChatGPT, click the sources or activity indicator. Both show the queries the model ran, not just the pages it read.
- Copy every query into the sheet, five runs per model. Tag each one generic (“best [category] for [segment]”, “[category] pricing comparison”) or brand-led (a specific company name inside the query).
- Note two things: which brands get named inside brand-led queries, and whether any brand that scored under 30% in the memory test still shows up in the final answer.
- Classify. Brand-led queries name only high-memory brands and low-memory brands never make the answer: memory-bound. Low-memory brands reach the answer through the generic queries: search-led. Most B2B categories land somewhere in between, and the ratio tells you how to split effort.
You have a memory score (appearance rate and mean rank) for you and two competitors on at least two models, a tagged list of real fan-out queries, and one word written at the top of the sheet: memory-bound or search-led.
- WATCH OUT
Model personalization is the trap here. ChatGPT with memory on, or a Google account that has visited your site a hundred times, will recall you far more often than a stranger's session would. Temporary chat, logged-out, clean profile, or the test measures your own browsing history.
- PRO TIP
Save the exact prompt text and re-run the memory test every quarter and after every major model release. Memory movement is the leading indicator that your category-authority work reached the training data. It will not show up month to month; it will show up in a step, when a new model ships.
- NO BUDGET?
The entire test above costs nothing and takes about an hour. The tracking platforms automate the repetition; they do not see anything you cannot see in the interface. Do it by hand first, at least once, because reading fifty raw fan-out queries teaches you your category's vocabulary in the model's own words.
The two-track plan for B2B SaaS
The test result decides the split, but almost every B2B SaaS needs both tracks, because the two run on different clocks. Track one pays this quarter and holds the line. Track two pays a model generation from now and is the durable position, since a remembered brand does not need to be re-found on every query.
| Track | The goal | The moves | How you know it is working |
|---|---|---|---|
| 1. Win the generic fan-out This quarter | Be in the pages retrieved for the 69% of queries that do not name a brand | Rank for the exact fan-out phrasings you copied from the test (they are ordinary search queries). Make those pages quotable. Get onto the third-party pages that already rank for them, per the off-page playbook. Put hard comparables where retrieval will find them: current ratings, review counts, real pricing. That is the override the experimental data measured. | You appear in answers on runs where the memory test says you were not recalled. Citation share moves before memory does. |
| 2. Get into memory This year, for next year's models | Move from under 30% appearance to the top five on the memory test | Consistent category association everywhere training corpora sample: the same identity sentence on your site, G2, Crunchbase, LinkedIn, docs. Co-occurrence with the top-5 brands in comparison content and roundups. Original data that gets re-cited under your name. Genuine community presence with the category term attached. | The quarterly memory test steps up after a model release. Brand-led fan-out queries start naming you. |
- Take the category vocabulary from your fan-out queries, the model's own words for what you are, and make it your identity sentence: “[Brand] is a [the model's category term] for [segment] that [outcome].” Put that sentence, verbatim, on your homepage, G2 listing, LinkedIn page, Crunchbase, and in your Organization schema. Models recall entities described consistently.
- Check your G2 category. If the model says “agency reporting software” and G2 has you filed under “business intelligence”, you are training the model to associate you with the wrong shelf. Request the category that matches the model's vocabulary.
- Build the co-occurrence surface: one honest comparison page against each of the top-5 memory brands from your test (“[Brand] vs [Top-5 brand]”), and an alternatives page that names all five. You are not trying to beat them on the page; you are teaching the corpus that you belong in the same sentence.
- Pitch inclusion in the specific roundups where the top-5 brands already appear together. Those pages are the highest-density category-association training data that exists, and the model has already told you which ones it trusts.
- Publish one named, original number this quarter (“the 2026 [category] benchmark”) with a permanent URL. Re-citations of a named stat are the cleanest brand-plus-category co-occurrence you can manufacture without paying anyone.
The identity sentence reads identically in five places, your G2 category matches the model's vocabulary, at least two comparison pages against memory-tier brands are live, and one pitch is out to a roundup that already lists the top five together.
- WATCH OUT
Do not buy your way into the co-occurrence surface. Paid placements on thin listicles sit in the third of citations that die at the next refresh, and a page that dies before the next training run contributes nothing to memory. Earned pages persist; rented ones evaporate. We made that argument in full in the rented-citations piece.
- PRO TIP
Expect a memory step, not a slope. Track two produces nothing visible for months and then a jump when the next model generation trains on the corpus you shaped. Tell leadership that timeline before you start, or the program gets cut in the flat part.
What this changes about measurement
Add one metric to the measurement system: memory rank, from the test above, tracked quarterly per model next to citation share. Citation share tells you whether you are winning the answer. Memory rank tells you whether you will keep winning it after the next model ships, and whether the brand-led queries are about to start working for you instead of against you. A rising citation share with a flat memory rank means track one is working and track two has not reached the training data yet. That is a normal and temporary state, and knowing that keeps you from misreading it as failure.
The honest part
Memory is the moat, and we have watched one form. When we made Swydo the most-cited brand in its category across every major model, the durable part was not any single page. It was that models started naming the brand without being prompted with it, in categories where it had not been in the first five a year earlier, with zero outreach emails and every citation earned. That is memory being built on the training clock, from corroboration the model could not ignore. It is the same order of operations we run as The Citation Engine, and it is exactly the two-track split above: win retrieval now, earn recall for the model your buyers use next.
