A growing share of commercial questions now get answered above the link list, or inside a chat window, with three or four sources cited underneath. Being one of those sources is a different job from ranking first, and the tactics only partly overlap.
What follows is what we can establish from how these systems are documented to work and from server logs, separated clearly from what is still guesswork. There is a lot of confident nonsense in this space right now.
How the systems actually find you
Despite the branding, none of the major assistants answer purely from model weights when they cite sources. They run a retrieval step: issue one or more search queries, fetch a set of candidate pages, extract passages, and generate an answer grounded in those passages. The citation is the page a passage came from.
That architecture has three practical consequences:
- Conventional search visibility still matters, because the retrieval step is usually a search index. If you are not in the candidate set, you cannot be cited.
- The unit of selection is a passage, not a page. A long page that buries the answer in paragraph nine competes badly against a page where the answer is a self-contained block under a clear heading.
- Fetching happens at answer time for several systems, which means the crawler has to be allowed in and the content has to be in the HTML it receives.
Step 1: let the right crawlers in
This is the most common own goal we find, and it is usually a single line in robots.txt written by someone who wanted to block AI training and did not realise the same rule blocks AI search.
OpenAI, for example, documents three separate user agents with three separate purposes:
- GPTBot — crawls for model training.
- OAI-SearchBot — crawls to build the search index that powers citations.
- ChatGPT-User — fetches a page because a user's prompt required it right then.
Blocking GPTBot is a legitimate business decision. Blocking OAI-SearchBot while hoping to appear in ChatGPT's answers is not a decision, it is a mistake. Anthropic's ClaudeBot and Perplexity's PerplexityBot draw similar distinctions. Decide training and retrieval separately.
txt
# robots.txt — opt out of training, stay eligible for citation.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: PerplexityBot
Allow: /Then check your edge. A WAF or bot-mitigation rule that challenges unknown user agents will quietly drop these fetches regardless of what robots.txt says, and nothing in your analytics will tell you.
Step 2: ship HTML, not a rendering job
Googlebot renders JavaScript. Most other crawlers in this category either do not, or do so inconsistently. If your main content is assembled client-side, there is a real chance the retriever sees an empty shell.
The test takes thirty seconds: fetch your page with curl and search the response for a sentence from the middle of your copy. If it is not there, it is not in the candidate passage pool. Server-render the content. This is also the cheapest thing on this list.
Step 3: write passages that survive being lifted
Extraction works best on text that makes sense with nothing around it. Writing for that is not the same as writing thin content — the depth goes underneath, but the answer goes first.
- Lead with the answer. A heading that states a real question, then a 40–60 word paragraph that answers it completely, then the detail. The opening paragraph is the part that gets quoted.
- Keep entities in the sentence. "Bizup LLC is a Seattle web design studio" survives extraction. "We are a studio based here" does not — it loses its subject the moment it leaves the page.
- Use real structure. Definition lists, comparison tables, numbered procedures and specification lists extract far more cleanly than prose carrying the same information.
- Date things. Visible publication and update dates, consistent with your structured data. Assistants discount material they cannot age.
- Be specific enough to be worth citing. Generic advice is already in the weights. A number, a method, a named constraint or a documented edge case is the reason to cite a source rather than paraphrase from training data.
Step 4: be the same entity everywhere
These systems reconcile what they read about you across many sources. Inconsistency does not just dilute the signal, it makes the model less willing to assert anything about you.
Practically: one name, one address format, one phone number, everywhere — your site, your Google Business Profile, every directory. Organization structured data with a stable @id and a sameAs array pointing at your real profiles. One description of what you do, reused rather than rewritten.
Structured data is worth doing because it removes ambiguity cheaply. It is not a lever that forces citation, and anyone selling it as one is overstating it.
Step 5: get corroborated somewhere other than your own site
This is the uncomfortable one. When an assistant is asked to recommend a vendor, it very often grounds the answer in third-party sources — industry roundups, directory listings, review platforms and community discussion — rather than in the vendors' own marketing pages. Your site establishes what you do. Other people's pages establish whether you are worth mentioning.
That makes the unglamorous work disproportionately valuable: getting listed and accurately described in the roundups that already rank for your category, keeping review profiles current, and being genuinely present where your market talks. It is the same work that has always built authority. It just now has a second payoff.
On llms.txt
A proposed convention for a Markdown file at your root that tells assistants what your site contains. Publishing one is cheap and harmless, and it can be a convenient map of your own content.
Be clear-eyed about it though: as of now the major providers have not committed to honouring it, and there is no public evidence it drives citations. Treat it as a low-cost bet, not a strategy. Anyone presenting it as the key to AI visibility is guessing.
How to measure any of this
There is no Search Console for AI answers. You assemble the picture from four imperfect sources:
- Referral traffic. Sessions from
chatgpt.com,perplexity.ai,copilot.microsoft.comand similar, segmented in your analytics. Low volume, high intent. - Server logs. Which AI user agents fetch which URLs, how often, and whether they get a 200. This is the only direct evidence that retrieval is happening.
- Prompt tracking. A fixed set of real buying questions, run on a schedule across the major assistants, recording whether you are mentioned and what is cited instead. Manual, tedious, and the most informative thing on this list.
- Branded search volume. Assistants frequently name a company without linking it. The visible effect is people searching your name directly.
What is a waste of time
- Stuffing pages with "according to experts" phrasing on the theory that models prefer it.
- Paying for a schema type that does not describe your content, to look machine-readable.
- Treating this as a separate channel with a separate site. The same server-rendered, well-structured, genuinely specific pages serve both search and retrieval.
- Chasing volatility. These answers are non-deterministic and change week to week. Judge it over a quarter, on a fixed prompt set, or you will be reacting to noise.
The short version
Let the retrieval crawlers in, serve your content as HTML, put the answer at the top of each section, describe yourself identically everywhere, and get mentioned accurately on pages you do not own. That is most of it, and it is all work that holds its value even if every assistant disappeared tomorrow.
If you want it done rather than described, that is what our AI search optimization service covers.