We spent eleven months tracking how six AI engines treated a single mid-sized software company, and the result contradicted the thing everyone assumes about generative search.
The assumption is that AI engines cite good content. Publish clearly, structure it well, answer the question, and the citations follow. That is the advice in almost every article written on the subject, and our data does not support it.
The company published consistently from month one. Citations stayed at roughly zero for the first two quarters. Then, over a period of a few months, they went from near-zero to 74 across all six engines. Nothing about the content changed in that window. What changed was something else entirely, and once you see it, a lot of confusing behaviour in AI search starts making sense.
What we tracked, and how
The subject was a warehouse management software provider selling to third-party logistics companies. Mid-sized, technical audience, competitive category with several well-funded incumbents.
We monitored citations across six surfaces: Google AI Overviews, Google AI Mode, ChatGPT, Perplexity, Microsoft Copilot and Gemini. Measurement came from a third-party tool rather than our own scorecard, which matters, because a citation count nobody can independently verify is a marketing asset rather than a measurement.
Alongside citations we tracked two other things: the volume and structure of published content, and the number of distinct websites linking to the domain.
| Metric | Start | End | Period |
| Citations across six AI engines | Near zero | 74 | 11 months |
| Referring domains | 113 | 357 | 11 months |
| Domain authority score | 19 | 55 | 11 months |
| Content published | Roughly 12 pieces per month, consistently | 11 months |
The finding: citations lagged authority, not content
Content output was steady from the first month. If citation eligibility were primarily a function of content quality and structure, citations should have risen roughly in step with publishing.
They did not. For the first two quarters, output climbed and citations did not move. The citation count only started rising after the number of distinct websites linking to the domain began climbing meaningfully.
Put plainly: in this dataset, the engines did not begin quoting the company until other websites had already started referencing it.
That ordering held across all six surfaces. We have not found a case in our own client data where it ran the other way, where citations arrived first and third-party references followed.
Why this is more intuitive than it sounds
A generative engine answering a question does two separate things. First it retrieves a set of candidate sources. Then it generates an answer from them, attributing as it goes.
Most published advice about optimising for AI search addresses the second stage. Write clear definitions. Use structural markers. Put the answer in the opening lines. Make passages standalone. All of that is real, and there is good research behind it, including a 2024 study out of Princeton that tested source-visibility levers across ten thousand queries and found that citations, quotations and statistics moved visibility by 30 to 40 percent.
But every one of those levers operates on content that has already been retrieved. If a domain never enters the candidate set, the quality of its formatting is irrelevant. It is competing in a round it was never entered into.
Retrieval appears to lean on the same broad trust signals that conventional search has always used, and third-party references are the oldest of those signals. That is the mechanism our data is most consistent with, and it is a boring answer to a question people want an exciting answer to.
The second finding: pages can be cited without being linked
The counterintuitive part came from looking at which individual pages earned visibility.
The strongest performing page on the site holds the top organic position for its target phrase and has zero websites linking to it directly. Not few. Zero. It is carried entirely by the trust the rest of the domain accumulated.
This is worth sitting with, because it cuts against the instinct that individual pages need individual links. In this dataset, authority behaved as a property of the domain, distributed across pages through internal structure rather than earned page by page.
The practical implication is that the site architecture matters more than most people assume. Content on this site was published into topical sections rather than a flat blog feed, with each section built around a coherent subject. Pages inside a well-built section inherited authority from the section. Pages that would have sat in an undifferentiated blog would not have.
What this changes about the advice
Three things follow, and the first one is uncomfortable.
Structure is necessary but not sufficient. Every piece of guidance about writing extractable, well-structured, clearly-defined content is correct and worth following. None of it will produce citations on a domain that has no third-party trust. Publishing more, better-structured content on a domain nobody references is optimising the wrong stage.
Measure both stages separately. The citation rate tells you what happened. It does not tell you why. Tracking the trust signals alongside the citation count is what lets you distinguish “we are not being retrieved” from “we are being retrieved and not chosen,” and those are different problems with different fixes.
Expect a lag. In our data the gap between authority moving and citations moving ran to about two quarters. Anyone promising visible movement in AI answers inside a few weeks is describing something they do not control.
The honest limits of this data
One company, eleven months, six engines. That is a case, not a study.
We have the aggregate citation figure across all six surfaces but we did not break out per-engine counts cleanly enough over the full period to publish them, and different engines almost certainly weight sources differently. Anyone claiming a universal ranking of what each engine wants is ahead of the available evidence, including us.
The clearest thing we can say is what the sequence looked like: references first, citations after. If your own tracking shows the reverse, that is genuinely interesting and worth publishing, because we have not seen it.
What we would test next
If we were designing this properly rather than observing a live client, the experiment would be a matched pair. Two domains, comparable authority, identical content strategy, with third-party references built aggressively on one and not the other. Track citation counts on both for twelve months.
Nobody has run that publicly, as far as we can find. Until somebody does, the field is working from cases like this one, and cases are worth publishing precisely because they are all anyone has.
We have written up the full method, including what we track per surface and what we deliberately do not measure, in our breakdown of how AI systems choose what to cite.
One thing worth doing this week
Whatever you conclude about the sequence, one check costs nothing and catches a surprising number of sites.
Confirm the crawlers can reach you. GPTBot, PerplexityBot, ClaudeBot and Google-Extended each need explicit access, and plenty of sites block one or more without realising, often through a robots.txt rule written years ago for a different reason. Google’s own guidance on AI features in search is the reference point for the Google surfaces.
A blocked crawler is the one failure mode where no amount of authority or structure helps at all. It is also the easiest thing on this list to fix.

