Model memory decides who gets searched

When ChatGPT or Gemini answers a buying question with web search on, it does not search for your category. It searches for what it already remembers about your category. The queries it fans out are written from memory, before a single page is retrieved.

We ran into this in client audits before we had a name for it. A B2B SaaS with quotable pages, present on the right review sites, technically clean, still absent from the answer. Not outranked. Never searched for. It lost two stages before retrieval started, and every fix we could have sold them would have applied to the wrong stage.

Our own 101-question benchmark across four engines measured which sources get cited, which is stage four of the four below. It could not explain the brands that never entered the running at all. Reading the fan-out queries did, and once we started running the memory test in this article on client categories, the pattern was consistent enough to plan around. Three independent lines of evidence now back it: Google documents the fan-out stage itself, public benchmarking sizes the effect at roughly 3×, and controlled experiments reproduce it with the brand name as the only variable.

how much more often a remembered brand gets searched, 55.7% vs 17.4%
67 / 39 / 17
search rate for top-5, rank 6 to 10, and not-in-memory brands
69%
of fan-out queries are generic; 31% name a brand, and most of those name a top-5 brand

What follows is the operator's version of it: the mechanism behind the numbers, how memory actually gets made and on what clock, an hour-long test that tells you whether your category is memory-bound, and the two-track plan that follows from the answer. Most coverage of this effect stops at “build category authority.” That is where this starts.

How a grounded AI answer is actually assembled

Four stages. Most AEO advice starts at the third.

  1. Recall. For a category it has seen enough of in training, the model can name a ranked set of brands without looking anything up.
  2. Fan-out. It writes the search queries it intends to run, and that recall shapes them.
  3. Retrieval. A search backend returns pages for those queries.
  4. Answer. The model writes from what came back, weighted by what it already believed.

Stage four is where citation tracking lives, and it is the only stage most teams measure. Stage two is where the shortlist is really set. Put them together and the puzzle resolves: you can be perfectly retrievable and never be retrieved, because the query that would have found you was never written.

[ THE LAYER ABOVE RETRIEVAL ]1. MEMORYBrands recalled beforetouching the webTop 5 in memory67%Rank 6 to 1039%Not in memory17%% OF TIMES THE BRAND GETS SEARCHED2. FAN-OUTThe queries the modeldecides to run69% generic“best [category] for [segment]”31% name a brand63% of those name a top-5 brand3. RETRIEVALPages that rankfor those queriesReddit, G2, listicles,docs, your siteThis is where mostAEO advice starts.Two stages too late.4. ANSWER1. Remembered2. Remembered3. Remembered4. You, only ifa generic queryfound you
Search rates by memory tier across 3,960 grounded answers in 9 industries, 2026. The model decides what to search before it searches.

What the data shows, and what it does not

The numbers in this piece come from the largest public run of this experiment we know of: 3,960 grounded answers across 66 US buyer questions and nine industries, five runs a day for twelve days, with 13,281 fan-out queries recorded. Memory was measured by asking a model to rank the top ten brands in a category before it touched the web. Search was measured by logging every query the model then fanned out. A brand counted as searched if it appeared in any of them.

Worth saying plainly: none of that is proprietary. The protocol is two prompts and a spreadsheet, which is why we can reproduce the core of it on any client category in an afternoon and why the test later in this article is the same experiment at small scale. Cite the big run for the magnitudes. Run your own for the only numbers that decide your budget, which are yours.

FindingThe numberWhat it means for you
Remembered brands get searched far more55.7% search rate when in the model's top 10, 17.4% when notRetrieval optimization has a ceiling set upstream of retrieval.
Rank inside memory mattersTop 5: 67%. Ranks 6 to 10: 39%. Not in memory: 17%Being “known” is not the goal. Being in the first five the model reaches for is.
Most fan-out is generic69% category queries, 31% brand-named; 63% of the named ones are top-5 brandsThe generic 69% is your way in if you are not remembered. They are ordinary search queries, and pages can rank for them.
It holds in every industry testedNot-remembered rates of 9 to 23%; remembered rates of 41 to 82% across nine verticals, including business softwareThis is not a consumer-brand quirk. B2B categories show the same shape.

Now the honest part, which that run states itself and which most people re-sharing the headline will skip. It is an association in exploratory data, not a proven cause. Brand prominence is the obvious confound: famous brands are both more likely to be remembered and more likely to be searched, so some of the effect is fame, not memory feeding search. It is two models over twelve days, US prompts only, and some industries rest on six to twelve prompts. Memory was measured on a different model than search, which makes the result harder to dismiss as one model's habit but also means the exact mechanism is not pinned down. The direction is robust. The magnitudes are a floor, not a law.

A single run is a single run, so we went looking for what else supports it and what argues with it.

What corroborates this, and what argues with it

Google documents the mechanism itself. Stage two is not an inference from watching a black box. When AI Mode launched, Google described it as using “a ‘query fan-out’ technique, issuing multiple related searches concurrently across subtopics and multiple data sources”. The model writes its own queries, and the vendor says so in plain language. That settles that the stage exists. The open question is only what shapes the queries.

Experimental work isolates the brand effect that observational data cannot. The strongest objection above is that fame drives both recall and searching, so you cannot separate them by watching. A June 2026 paper by Chu and Hou attacks it the other way, by holding product specifications constant and swapping only the brand name. Across 670 valid trials where the options were identical on paper, the real brand was recommended 100% of the time. That is an experiment, not a correlation, and it isolates precisely the variable the big observational run admits it cannot.

Two more results from that paper matter more for planning than the headline does. When genuine quality information was available, product attributes explained 82.4% of the variance in what got recommended and brand identity only 1.2%. And under retrieval-augmented generation, the incumbent advantage largely collapsed. Read those together and the two-track plan later in this piece falls out of the evidence rather than out of our opinion: memory decides when the model has nothing else to go on, and evidence in the retrieved context overrides it.

The size of the override is the practical gift. In that study the brand advantage broke with a 0.075-star rating edge, 1.6 times the review count, or a 7.3% lower price. For a B2B SaaS that is not an abstraction. It means a visible, current, specific G2 profile is not a nice-to-have sitting next to your AEO work. It is the cheapest known way to beat a better-remembered competitor at the moment of comparison.

The mentions signal points the same direction. Ahrefs' 75,000-brand study found brand web mentions correlate with AI Overview visibility at 0.664 versus 0.218 for backlinks. If recall is built from corroborated mentions, that is the fingerprint you would expect to find, and it is measured on a different surface by a different party.

And the honest counterweight. The literature is not unanimous. Lichtenberg, Buchholz and Schwöbel found that an LLM recommender showed less popularity bias than conventional recommender systems on movie recommendations, with no mitigation applied. Different domain, no retrieval step, consumer catalogue rather than B2B category. But it is a real result pointing the other way, and anyone telling you models simply favour the famous in all conditions has not read it.

What survives all of it is narrower than the headline and more useful. The fan-out stage is documented. A known-brand advantage at that stage is measurable observationally and reproducible experimentally. Retrieved evidence beats it. That is enough to plan against, and it is the whole argument for running both tracks instead of picking one.

You can be perfectly retrievable and never be retrieved. The query that would have found you was written from a memory you are not in.

The part nobody explains, how memory gets made

A model's memory of your category is a statistical residue of its training data: how often your brand appeared, how consistently it was described, and how tightly it co-occurred with the category term and with the other brands the model already associates with that category. Three properties of that process matter more than any tactic.

  • It is built from corroboration, not from your site. One page saying you are a “revenue intelligence platform” is a claim. Fifty independent pages saying it, in reviews, roundups, community threads, docs and press, is a fact the model can hold. This is the same signal behind Ahrefs' finding that brand mentions correlate with AI visibility at 0.664 versus 0.218 for backlinks. Mentions are training data. Links mostly are not.
  • It is built from co-occurrence. Watch a ride-hailing fan-out and you see Waymo searched alongside Uber and Lyft, not because anyone asked about it but because the model files it on the same shelf. A brand that keeps appearing next to the top five in comparison content inherits some of their category association. Proximity is a signal.
  • It updates on the training clock, not yours. Memory is frozen at a model's training cutoff and refreshes when the next generation ships, which in 2026 means a lag of months between earning a mention and the model recalling it. Everything you do for memory today is an investment in the model your buyers will be using next year. Search-led wins are the bridge until then.

One inference is ours, drawn from the client categories we have tested rather than from anyone's published run. Business software categories are narrow, and narrow categories have thinner memory: fewer brands the model recalls with confidence, and a lower bar to get into the top five. The public data points the same way, since its clearest example of an unremembered brand still getting searched came from a specialized software category. Read that two ways at once. B2B SaaS is more search-led than consumer categories, so live content can still win this quarter, and the memory slots are cheaper to take now than they will ever be again.

Which regime is your category in? Test it in an hour

There are two regimes. In a memory-bound category, the brand-named queries only ever name remembered brands and an unremembered brand almost never surfaces. In a search-led category, the generic queries still pull in brands the model did not recall. You need to know which one you are in before you spend a dollar, and you can find out today with nothing but the chat interfaces.

[ RUN THE MEMORY TEST, EXACTLY ]
  • Open a temporary chat in ChatGPT (the memory feature off, search off), a fresh Gemini chat, and Claude. Use a browser profile that has never discussed your brand; the personalization features will otherwise flatter you.
  • Paste this, filling in your category and segment: “Without searching the web, list the 10 [category] tools you would consider for a [segment] company. Rank them by how confident you are, most confident first. Answer only from what you already know.”
  • Run it 10 times per model in fresh chats. In a sheet, record for each run whether your brand appeared and at what rank. Do the same for your two closest competitors while you are there.
  • Compute two numbers per model: appearance rate (runs where you were named, out of 10) and mean rank when named. Appearance rate under 30% means you are effectively not in memory. Mean rank 1 to 5 with high appearance is the top tier that gets searched about 67% of the time.
[ RUN THE FAN-OUT TEST, EXACTLY ]
  • Same question, search on. In Gemini, expand the “Searched for” or sources panel under the answer; in ChatGPT, click the sources or activity indicator. Both show the queries the model ran, not just the pages it read.
  • Copy every query into the sheet, five runs per model. Tag each one generic (“best [category] for [segment]”, “[category] pricing comparison”) or brand-led (a specific company name inside the query).
  • Note two things: which brands get named inside brand-led queries, and whether any brand that scored under 30% in the memory test still shows up in the final answer.
  • Classify. Brand-led queries name only high-memory brands and low-memory brands never make the answer: memory-bound. Low-memory brands reach the answer through the generic queries: search-led. Most B2B categories land somewhere in between, and the ratio tells you how to split effort.
YOU DID IT RIGHT IF

You have a memory score (appearance rate and mean rank) for you and two competitors on at least two models, a tagged list of real fan-out queries, and one word written at the top of the sheet: memory-bound or search-led.

  • WATCH OUT

    Model personalization is the trap here. ChatGPT with memory on, or a Google account that has visited your site a hundred times, will recall you far more often than a stranger's session would. Temporary chat, logged-out, clean profile, or the test measures your own browsing history.

  • PRO TIP

    Save the exact prompt text and re-run the memory test every quarter and after every major model release. Memory movement is the leading indicator that your category-authority work reached the training data. It will not show up month to month; it will show up in a step, when a new model ships.

  • NO BUDGET?

    The entire test above costs nothing and takes about an hour. The tracking platforms automate the repetition; they do not see anything you cannot see in the interface. Do it by hand first, at least once, because reading fifty raw fan-out queries teaches you your category's vocabulary in the model's own words.

The two-track plan for B2B SaaS

The test result decides the split, but almost every B2B SaaS needs both tracks, because the two run on different clocks. Track one pays this quarter and holds the line. Track two pays a model generation from now and is the durable position, since a remembered brand does not need to be re-found on every query.

TrackThe goalThe movesHow you know it is working
1. Win the generic fan-out
This quarter
Be in the pages retrieved for the 69% of queries that do not name a brandRank for the exact fan-out phrasings you copied from the test (they are ordinary search queries). Make those pages quotable. Get onto the third-party pages that already rank for them, per the off-page playbook. Put hard comparables where retrieval will find them: current ratings, review counts, real pricing. That is the override the experimental data measured.You appear in answers on runs where the memory test says you were not recalled. Citation share moves before memory does.
2. Get into memory
This year, for next year's models
Move from under 30% appearance to the top five on the memory testConsistent category association everywhere training corpora sample: the same identity sentence on your site, G2, Crunchbase, LinkedIn, docs. Co-occurrence with the top-5 brands in comparison content and roundups. Original data that gets re-cited under your name. Genuine community presence with the category term attached.The quarterly memory test steps up after a model release. Brand-led fan-out queries start naming you.
[ START TRACK TWO, EXACTLY ]
  • Take the category vocabulary from your fan-out queries, the model's own words for what you are, and make it your identity sentence: “[Brand] is a [the model's category term] for [segment] that [outcome].” Put that sentence, verbatim, on your homepage, G2 listing, LinkedIn page, Crunchbase, and in your Organization schema. Models recall entities described consistently.
  • Check your G2 category. If the model says “agency reporting software” and G2 has you filed under “business intelligence”, you are training the model to associate you with the wrong shelf. Request the category that matches the model's vocabulary.
  • Build the co-occurrence surface: one honest comparison page against each of the top-5 memory brands from your test (“[Brand] vs [Top-5 brand]”), and an alternatives page that names all five. You are not trying to beat them on the page; you are teaching the corpus that you belong in the same sentence.
  • Pitch inclusion in the specific roundups where the top-5 brands already appear together. Those pages are the highest-density category-association training data that exists, and the model has already told you which ones it trusts.
  • Publish one named, original number this quarter (“the 2026 [category] benchmark”) with a permanent URL. Re-citations of a named stat are the cleanest brand-plus-category co-occurrence you can manufacture without paying anyone.
YOU DID IT RIGHT IF

The identity sentence reads identically in five places, your G2 category matches the model's vocabulary, at least two comparison pages against memory-tier brands are live, and one pitch is out to a roundup that already lists the top five together.

  • WATCH OUT

    Do not buy your way into the co-occurrence surface. Paid placements on thin listicles sit in the third of citations that die at the next refresh, and a page that dies before the next training run contributes nothing to memory. Earned pages persist; rented ones evaporate. We made that argument in full in the rented-citations piece.

  • PRO TIP

    Expect a memory step, not a slope. Track two produces nothing visible for months and then a jump when the next model generation trains on the corpus you shaped. Tell leadership that timeline before you start, or the program gets cut in the flat part.

What this changes about measurement

Add one metric to the measurement system: memory rank, from the test above, tracked quarterly per model next to citation share. Citation share tells you whether you are winning the answer. Memory rank tells you whether you will keep winning it after the next model ships, and whether the brand-led queries are about to start working for you instead of against you. A rising citation share with a flat memory rank means track one is working and track two has not reached the training data yet. That is a normal and temporary state, and knowing that keeps you from misreading it as failure.

The honest part

Memory is the moat, and we have watched one form. When we made Swydo the most-cited brand in its category across every major model, the durable part was not any single page. It was that models started naming the brand without being prompted with it, in categories where it had not been in the first five a year earlier, with zero outreach emails and every citation earned. That is memory being built on the training clock, from corroboration the model could not ignore. It is the same order of operations we run as The Citation Engine, and it is exactly the two-track split above: win retrieval now, earn recall for the model your buyers use next.

DO THIS NEXT

Want to run both tracks yourself with checkpoints? The free SEO + AEO Academy covers the memory test inside prompt mapping and the engine audit, and track two in trusted-source engineering. Rather have it done for you? Book a strategy call and we'll map your category live, for free, on the call.

Frequently asked

What is model memory in AI search?
The set of brands a language model can recall for a category from its training data, before it searches the web, ranked by how confidently it recalls them. Google documents this fan-out stage in AI Mode, where the model issues multiple related searches it writes itself. Because those queries are written from recall, brands the model remembers get searched far more often: about 56% of the time versus 17% for brands it does not recall, across a 2026 benchmark of 3,960 grounded answers.
Why does ChatGPT search for some brands and not others?
Because the queries it fans out are written before retrieval, from what it already remembers about the category. About 31% of fan-out queries name a specific brand, and most of those name brands in the model's top five for the category. The remaining 69% are generic category queries, which is the route in for brands the model does not yet recall.
How do I get my brand into an LLM's memory?
Memory is a residue of training data, so the work is category association across the sources models train on: one consistent identity sentence everywhere your brand is described, a G2 category that matches the model's own vocabulary, comparison and alternatives content that places you next to the brands the model already recalls, inclusion in the roundups where those brands appear together, and original data that gets re-cited under your name. It shows up when the next model generation trains, not next month.
Does model memory matter if the model searches the web anyway?
Yes, because memory shapes what it searches for. A brand outside memory can still reach the answer through generic category queries, which is why quotable pages and presence on retrieved sources still pay this quarter, but it has to be re-found on every query. A remembered brand gets searched for by name and does not depend on being found. Experimental work supports both halves: brand identity dominates recommendations when options look identical, and retrieval-augmented generation largely collapses that advantage once real comparables are in the context. Search-led wins are the bridge; memory is the durable position.
How long does it take to change what an AI model remembers about a brand?
A model generation. Memory is frozen at a model's training cutoff and refreshes when the next version ships, so mentions earned today typically surface in recall months later, as a step change rather than a gradual slope. Track it by re-running a fixed memory test quarterly and after every major model release, and set that expectation with leadership before starting.
How can I test whether an AI model knows my brand?
With search disabled and personalization off, ask the model to list the ten tools it would consider in your category ranked by confidence, ten times in fresh chats, and record how often you appear and at what rank. Under 30% appearance means you are effectively not in memory. Then run the same question with search on and read the queries it fanned out to see whether the category is memory-bound or still reachable through generic searches.
Jay Kang, founder of Enginekick
Written by Jay Kang // Founder, Enginekick

Jay is a B2B SaaS organic-growth operator with 10+ years across technical SEO, content and AI search. He made Swydo the most-cited brand in its category across every major LLM with zero outreach, grew AgencyAnalytics from ~$9M to $20M+ ARR, built Enginekick OS for the agency's operating system, and builds Pageoptimized, the SEO and AI-visibility platform the work runs on.

Connect on LinkedIn

Be searched for
by name.

Book a strategy call. We'll run the memory and fan-out tests on your category live, across every major model, and show you which track to start first.

Book a strategy call