Site speed used to be argued as a user-experience issue with a ranking tiebreaker attached. That argument still holds, and it is now the smaller half of the case.
The larger half is that an increasing number of the systems deciding whether to show your content fetch it programmatically, on a deadline, with no patience and no human to wait. For those systems, your server's behaviour is not a quality signal — it is the difference between being read and not existing.
Three gates between your content and an answer
A page has to pass all three. Failing any one of them is indistinguishable, from the outside, from having nothing to say.
- It has to be allowed. robots.txt, your edge rules and your bot mitigation all get a vote before anything is fetched.
- It has to be served quickly. Automated fetchers operate with timeouts. A slow origin is a dropped fetch, and nothing in your analytics will record the loss.
- It has to contain the content in the response. Not after hydration. In the bytes the fetcher receives.
Gate one: who you are actually allowing
The most common own goal we find is a robots.txt written by someone who wanted to opt out of AI training and did not realise the same rule blocks AI search. These are separate decisions and the agents are separate:
- GPTBot — OpenAI's crawler for model training.
- OAI-SearchBot — OpenAI's crawler for the search index that produces citations. Blocking this one while hoping to appear in ChatGPT's answers is not a trade-off, it is a mistake.
- ChatGPT-User — fetches a page because a user's prompt required it at that moment.
- PerplexityBot — Perplexity's indexing crawler, documented alongside a separate user-triggered agent.
- Google-Extended — not a crawler at all. It is a robots.txt token controlling whether content Googlebot already fetched may be used for Gemini and Vertex AI generative products. Google documents that it does not affect inclusion or ranking in Google Search.
Anthropic's ClaudeBot and the other major providers draw similar distinctions. Check each provider's current published documentation before you write the file — the names and the semantics change, and a rule written from a blog post two years old is likely to be wrong in both directions.
Gate one and a half: your edge has opinions you did not express
This is the part that catches technically careful teams. robots.txt is a request to well-behaved crawlers. Your CDN, WAF and bot-mitigation layer are enforcement, and they frequently challenge or block user agents they do not recognise — which is, by definition, every new one.
Cloudflare, for instance, ships a one-click AI-crawler block and has made blocking the default for newly onboarded domains. That may be exactly what you want. What you do not want is for it to be true without anyone having decided it, which is the usual situation.
There is no way to discover this from the outside. You have to look at your own logs, filtered by user agent, and check the status codes — then look at your edge dashboard and read the rules that are actually enabled.
bash
# What do the AI fetchers actually get from you?
grep -Ei 'GPTBot|OAI-SearchBot|ChatGPT-User|PerplexityBot|ClaudeBot' access.log \
| awk '{print $9}' | sort | uniq -c | sort -rn
# Any 403 or 429 in that list is your edge answering on your behalf.Gate two: timeouts are real
A human visitor on a slow page waits, sighs and sometimes stays. An automated fetcher does not. It has a deadline, and when the deadline passes it moves to the next candidate — which is a competitor's page.
Time to first byte is the number that governs this, and it is the one front-end optimisation cannot help with. It is origin configuration, caching strategy, and whatever database query runs before the first byte leaves. A statically generated page served from cache answers in single-digit milliseconds; the same content assembled per-request by a plugin stack on shared hosting can take over a second before anything is sent.
There is a second-order effect too. Crawl budget is partly a function of how fast you respond: a fast origin gets more of its pages fetched in the same window. For a large site that is the difference between your whole catalogue being known and a fraction of it being known.
Gate three: ship HTML
Googlebot renders JavaScript. Most other crawlers in this category either do not or do so inconsistently, and none of them are obliged to tell you. If your main content is assembled client-side, there is a real chance the retriever receives a shell with a loading spinner in it.
The test takes thirty seconds and it is not optional:
bash
# Is your content in the response, or assembled later?
curl -s https://example.com/your-key-page \
| grep -c "a distinctive sentence from the middle of the page"
# 0 means the content is not there. Nothing downstream can fix that.Server-render the content. On a modern framework this is the default rather than an achievement, and it is the cheapest item on this entire list.
Why the two concerns converged
For years, performance and discoverability were argued separately — one to the design team, one to the marketing team. The merge is not philosophical, it is mechanical: the same three properties decide both.
A server-rendered page is readable by a person with a flaky connection and by a fetcher that does not run JavaScript. A fast origin serves an impatient human and a crawler with a timeout. Clean semantic structure makes a page navigable by a screen reader and extractable as a passage. In each pair, the second benefit arrived later and cost nothing extra.
Which is why we do not sell these as separate projects when they overlap. The speed work and the AI search work share most of their foundation; what differs is what you measure afterwards.
The audit, in order
- Read your own robots.txt and decide, deliberately, training versus retrieval for each provider.
- Read your edge rules. Find out what is being blocked or challenged without your knowledge.
- Check the status codes the AI user agents actually receive in your logs.
- Measure TTFB on your templates, uncached, from more than one region.
- curl your three most important pages and confirm the content is in the response.
Five checks, half a day, and no new content required. It is the least glamorous work in this category and the only part where the failure mode is total rather than gradual. For what to do once the gates are open, see how to get cited in AI search.