Research 6 min read

Where models actually look

Ask a model about a brand and you get a confident paragraph. Ask where that paragraph came from and the honest answer is that it came from several places at once, weighted in ways nobody outside the lab can see precisely.

You can still say useful things about it, because the pattern is visible in the citations, and because it has been fairly stable for long enough to plan around.

Two different mechanisms, often confused

The first is what the model absorbed during training. This is where the general shape of your brand lives — what category you are in, roughly what you cost, the adjectives that follow your name. It is old, it is slow to change, and no amount of publishing this quarter will shift it this quarter.

The second is what the model fetched just now, mid-answer, to check itself. This is where current facts come from: your pricing page, a recent review, whatever ranked for the search it ran behind the scenes.

Most confusion about AI visibility comes from treating these as one system. They respond to completely different work on completely different timescales.

Signal Shapes Moves in Where effort goes
Training data Category, positioning, tone Model generations Being written about, widely
Retrieved pages Facts, prices, features Days to weeks Pages that state things plainly
Third-party reviews Recommendations, caveats Weeks to months Presence where reviews live
Your own site Corrections, specifics Days Structure and clarity

Third-party text does most of the work

The single most consistent thing across the answers I have read: models lean on other people's writing about you more heavily than on yours.

That makes sense. If the question is which option is best, a page belonging to one of the options is a source with an obvious interest. A roundup, a review, a forum thread where somebody with no stake describes the trade-offs — those read as evidence in a way marketing copy does not.

The uncomfortable implication is that your most valuable brand asset may be a page you don't own, can't edit, and didn't commission.

This is why the brands that do well in AI answers are frequently not the ones with the best websites. They are the ones that have been discussed a lot, in places where discussion accumulates.

Plain statements beat persuasive ones

Within your own site, the pages that get cited share a quality that has nothing to do with design or conversion rate. They state facts in complete sentences, near the top, without conditions.

"Plans start at nine pounds a month" is liftable. "Flexible pricing designed around your needs" is not liftable, because there is nothing in it to lift. A model asked what something costs will pass over the second sentence and go looking for the first, and if the first does not exist on your site it will find it somewhere else — a review, a comparison page, a competitor's summary of you, with whatever inaccuracies that carries.

Documentation and support pages often outperform product pages for exactly this reason. They were written to answer a question rather than to close a sale, so they answer questions.

Absence compounds quietly

The failure mode that does the most damage is not being described badly. It is not being present in the material at all.

If nobody has written a comparison that includes you, the model assembling a comparison has nothing of yours to reach for, and it will build the comparison from the brands that do appear. You are not beaten in that answer; you are simply not in it, and no one involved notices.

This is the argument for tracking rather than auditing once. A single check tells you where you stand today. Watching the same questions over months tells you whether you are entering the conversation or slowly falling out of it — and falling out is gradual enough that it is invisible without a baseline.

What this suggests about effort

Three things follow, none of them quick.

Put facts on your own pages in plain language, high up, and keep them current. This is the cheapest work available and it is genuinely effective for the retrieval half of the system.

Be present where third parties write. Reviews, comparisons, communities where your category is discussed. Slow, unglamorous, and it does more for the training half than anything you publish yourself.

And track the questions rather than the pages. What you need to know is whether you are in the answer, and that is not a fact about any page you own.

Why the two mechanisms drift apart

A brand can be current and wrong at the same time. The retrieval half reads today's pricing page and reports it correctly, while the training half still carries a description of you from two positioning changes ago.

The answer that results is internally inconsistent in a way no human writer would produce: an accurate price attached to an outdated characterisation, or a current feature list introduced by a sentence describing the company you stopped being some time ago.

There is no single fix for this, because the two halves are not fed by the same pipe. The retrieval half you can address this month. The training half you address by giving the next generation of models something better to read, which means the work you do now pays out on a schedule you do not control.

Reading citations as a map of the gap

The practical diagnostic is to compare what an answer claims against what it cites.

Where the claim is supported by a citation, you are looking at retrieval, and the source is right there to be checked. Where the claim has no citation attached — the confident summarising sentence, the characterisation, the adjective — you are usually looking at absorbed knowledge, and the source is diffuse and historical.

Uncited claims are the ones worth reading most carefully, because they are the ones you cannot trace and cannot quickly correct. They are also, in most answers, the sentences that do the actual persuading.

The parts of this you can control

It is worth being blunt about the split, because effort spent on the wrong half feels productive and achieves very little.

You control what your own pages say and how legibly they say it. You control, partially and slowly, how much third-party material exists about you and what it contains. You do not control how any of it is weighted, whether a given answer retrieves anything at all, or how a future model generation will summarise the category.

That leaves a narrower set of actions than most marketing channels offer, and it makes measurement more important rather than less. When you cannot steer directly, knowing which way you are drifting is most of the job.

Keep reading

← All articles

How visible is your brand in AI answers?

Tell us your domain and we'll send you a visibility report — free, no card.