Skip to content
NEWAI search optimizationSee how

AI search optimization · GEO / AEO

Get cited when someone asks an AI about your market

A growing share of research now ends inside an answer rather than on a results page. The buyer asks ChatGPT, Perplexity or Google’s AI Overview which firm to use, and three sources get named. This is the work that makes you one of them — and the honest version of what is known about it, which is less than most agencies imply.

Surfaces
AI Overviews, ChatGPT, Perplexity, Copilot
Also called
GEO · AEO · LLMO
Prerequisite
A crawlable, sound site
Guarantees
None — and nobody can offer one

The problem

Rankings holding steady, traffic sliding

The pattern is now familiar. Positions in Search Console look stable, impressions may even be up, and clicks are down. The answer is being assembled above the results from a handful of sources, and the user never scrolls.

Being in the organic top ten no longer guarantees you are one of the sources the answer is built from. Those are selected for different reasons: whether a model can identify you as a distinct entity, whether your pages are machine-extractable, whether your claims are specific enough to quote, and whether other sources corroborate you.

That is a different job from classic SEO, though it sits on the same foundation. It is also genuinely new, which means a lot of confident advice about it is invented. We will tell you which parts are established, which are reasonable inference, and which are nobody’s idea of proven.

What this looks like in your data

  • Impressions flat or rising while clicks fall — classic zero-click pressure
  • Competitors named in AI answers for queries you rank well for
  • Ask an assistant about your category and your brand does not appear
  • Ask it about your brand directly and it gets facts wrong, or confuses you with another company
  • Referral traffic appearing from assistant domains, but barely any
  • Nobody on your team can say whether any of this is getting better or worse

For reference

What we can measure — and what we cannot

Honest measurement is the hardest part of this discipline. Here is exactly what each signal does and does not tell you.

We report all six. Any agency reporting only the flattering one is managing your perception rather than your visibility.
SignalHow we capture itWhat it does not tell you
Citation rate across a prompt panelA fixed set of buyer prompts re-run on a schedule across the major assistants, logged with datesNothing about volume. Appearing in an answer does not mean anyone asked that question
Referral traffic from assistant domainsAnalytics segments for the assistant referrers that pass oneMost assistant sessions pass no referrer or get grouped as direct, so this undercounts, often badly
AI crawler hits in server logsLog filtering by user agent, with verification against published IP rangesThat you were retrieved or cited — only that you were fetched
Branded search volumeSearch Console and keyword tools, tracked as a trendAttribution. A rise is consistent with assistant-driven discovery but does not prove it
Structured data and rendering healthValidation and raw-HTML checks on every deployOutcomes. These are inputs that remove obstacles, not results
Accuracy of what assistants say about youReading the actual text returned about your brand, not just whether you appearHow many buyers saw it. But a wrong fact repeated is worth fixing regardless

What is included

What we actually do

Six workstreams. The first three are foundations that also help classic search; the last three are specific to being used as a source.

Entity foundation

Who you are, unambiguously
  • A single Organization node with a stable @id referenced by every other schema object on the site
  • sameAs links to every profile that corroborates you — LinkedIn, Business Profile, Crunchbase, trade bodies
  • Name, address and phone made identical everywhere they appear, on and off site
  • Disambiguation from similarly named companies, which is a common cause of wrong AI answers
  • Named authors with real credentials and author pages that are themselves entities

Structured data as a graph

Not isolated blobs
  • Organization, Service, FAQPage, BreadcrumbList, Article and Product as appropriate
  • Nodes linked by @id rather than repeated — one source of truth per fact
  • author, datePublished and dateModified populated honestly on every article
  • Schema validated and kept in sync with the visible page, because contradicting your own markup is worse than omitting it
  • Review and rating markup used only where it is genuinely eligible

Machine-readable pages

Extractability
  • Server-rendered HTML — content that only appears after client-side JavaScript may never be seen
  • One idea per section, with a heading that states it
  • Answer-first paragraphs: the claim in the first sentence, support after
  • Real lists, real tables and real definitions instead of text baked into images
  • Clear, self-contained passages, since retrieval operates on chunks rather than whole pages

Crawler access policy

Deliberate, not accidental
  • A robots.txt policy covering GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended, Bingbot and CCBot
  • The trade-offs explained per crawler before you decide, not after
  • Verification that your CDN or firewall is not silently blocking the ones you want
  • An /llms.txt manifest pointing to your canonical reference pages
  • Server-log checks for which AI crawlers are actually visiting

Citation-worthy content

The part most skip
  • Original material only you can publish — your own data, process, pricing logic, methodology
  • Direct definitional sentences for the terms in your category
  • Comparison tables, including honest ones where you are not the answer
  • Specific figures with dates and sources, because vague claims are not quotable
  • Pages that answer the question a buyer actually types, not the keyword a tool returned

Corroboration and measurement

On-site is half the job
  • Presence on the third-party sources assistants lean on — directories, roundups, industry press, community threads
  • A tracked panel of real prompts, re-run on a schedule across the major assistants
  • Referral traffic segmented by assistant domain in analytics
  • Branded search volume watched as a proxy for assistant-driven discovery
  • A monthly readout of where you are named, where you are not, and what changed

The technical detail

What actually makes a page quotable

Retrieval-augmented systems do not read your site the way a person does. They fetch, chunk, embed, retrieve and then summarise with attribution. Each stage can lose you.

  1. Rendering: if it needs JavaScript, assume it is invisible

    Classic search engines render JavaScript, with delay. Many AI crawlers do not render at all — they take the raw HTML response and move on. Content injected on the client, tabs that fetch on click, and accordions whose content is loaded lazily are all at risk of simply not existing.

    • Server-render or statically generate the content that matters
    • Check the raw HTML response rather than the inspector’s rendered DOM
    • Keep critical content out of iframes, canvas and images of text
    • Make sure a 403 from your CDN is not what the crawler is receiving
  2. Chunking: write sections that survive being cut out

    Documents are split into passages before retrieval. A passage that only makes sense after three paragraphs of build-up loses its meaning when extracted. Self-contained sections retrieve better.

    • State the claim in the section’s first sentence
    • Repeat the subject rather than relying on 'it' and 'this' across headings
    • Use headings that are statements or questions, not one-word labels
    • Keep one topic per section; split rather than nest
  3. Entity clarity: be one unambiguous thing

    Models reason over entities, not strings. If your company name is spelled three ways, your address differs between your site and your Business Profile, and nothing links your profiles together, the system has several weak partial entities instead of one strong one.

    • One canonical legal and trading name, used consistently
    • A single Organization node with a stable @id, referenced everywhere
    • sameAs pointing at every profile that independently confirms you exist
    • Author entities with credentials, not 'admin' bylines
  4. Structured data: a connected graph, not scattered tags

    Schema does not make a model like you. It removes ambiguity about what a page is and what the facts on it are, which is a different and more durable benefit. The win comes from a connected graph whose nodes reference one another and whose values match the visible page exactly.

    • Link nodes with @id instead of duplicating organisation data per page
    • Keep FAQPage markup to questions genuinely answered on the page
    • Populate dateModified truthfully — recency is a real retrieval signal
    • Validate on every deploy; broken markup is silently ignored
  5. llms.txt: cheap, useful, and not the lever

    llms.txt is a proposed convention — a Markdown file at your root that points to your canonical documentation in a clean, link-rich form. It is trivial to ship and does no harm. What it is not, today, is a confirmed ranking or citation input for the major assistants. We ship it because it costs an hour and may matter later. We will not sell it to you as the strategy, and an agency leading with it is telling you something about their depth.

    • Publish /llms.txt listing your canonical reference pages with one-line descriptions
    • Keep it generated from the same source as your sitemap so it cannot go stale
    • Treat it as housekeeping, not as the programme
  6. Crawler policy: decide it on purpose

    There is a real trade-off and it deserves five minutes of your attention rather than a default. Blocking training crawlers protects your content from being absorbed into model weights; it can also reduce the chance of being represented at all. The controls are not one switch.

    • GPTBot is OpenAI’s training crawler; OAI-SearchBot serves its search and browsing features — they are separate decisions
    • Google-Extended governs Gemini training, and does not remove you from Google Search or, in general, from AI Overviews
    • PerplexityBot and ClaudeBot are separate agents with their own directives
    • CCBot feeds Common Crawl, which many downstream datasets are built from
    • Whatever you choose, verify it in your server logs rather than trusting the file
  7. Corroboration: the part that is not on your website

    Assistants synthesise across sources. A claim that appears only on your own site is weaker than one that also appears on an industry list, a review platform, a trade publication and a community thread. This is the same logic as links, applied to mentions, and it is the slowest part of the work.

    • Accurate, complete profiles on the platforms your category is listed in
    • Inclusion in genuine roundups and comparisons, earned not bought
    • Participation where your buyers actually discuss the problem
    • Consistency of facts across all of it — contradictions reduce confidence
  8. Measurement: trends over screenshots

    Assistant outputs are non-deterministic and personalised, so a single favourable response proves very little. Measurement has to be repeated sampling of a fixed prompt set, read as a trend, combined with the server-side signals you can actually count.

    • A fixed prompt panel, re-run on a schedule, results logged with dates
    • Referral traffic from assistant domains, segmented in analytics
    • Branded search volume as a proxy for discovery you cannot attribute directly
    • Server logs for AI crawler visits — frequency, depth and status codes

Our process

How we approach a new engine

The methodology is deliberately empirical. We measure a baseline, change one category of thing at a time, and re-measure.

  1. Understand needs

    Baseline what the assistants say now

    We build a panel of the prompts your buyers would plausibly use, run them across the major assistants, and record what is returned: who is named, what is cited, and what is said about you. Alongside that, a technical read of entity signals, schema, rendering and crawler access.

    You getPrompt baseline and technical entity audit

  2. Strategize

    Decide where you can plausibly win

    Some answers are locked up by publishers and marketplaces and are not worth chasing. We pick the question space where a specialist firm is a credible source, and the corroboration gaps worth closing.

    You getTarget prompt set and prioritised plan

  3. Create & build

    Ship foundations, then content

    Entity and schema graph, rendering and crawler policy first, because content cannot be cited if it cannot be read. Then the pages designed to be quoted.

    You getImplemented foundations and published pages

  4. Optimize & grow

    Re-run the panel, report honestly

    The same prompts, re-run monthly. Assistant outputs vary between runs, so we read trends across repeated samples rather than celebrating a single good screenshot.

    You getMonthly citation report with trend, not anecdote

Honest scoping

Who this is genuinely for

Worth doing if

  • Your buyers research before they buy — considered purchases, professional services, B2B
  • You already have decent organic visibility and want to protect it as behaviour shifts
  • You have real expertise or proprietary data that could be worth citing
  • You can publish; this work needs material, not just markup
  • You can judge a programme on leading indicators for a few months before it shows in revenue

Not yet if

  • Your site has unresolved crawling, indexing or rendering problems — fix those first with SEO, then come back
  • You want a guarantee of citation. No vendor can deliver one; these systems re-rank on every query and vary between users
  • You are unwilling to publish anything new
  • Your buyers find you through referral and walk-ins, and search is not part of the journey
  • You want this instead of SEO rather than on top of it

One more thing. This is a young discipline. We would rather be the agency that tells you which parts are evidence and which are hypothesis than the one selling certainty about a system nobody outside the labs can see inside.

Questions

Before you ask us

The things people ask about AI Search Optimization before they get in touch. If yours is not here, ask directly — you will get a straight answer rather than a brochure.

Ask a question

Is this just SEO with a new name?

It shares a foundation and diverges at the top. Both need a crawlable, fast, well-structured site, and both reward genuine expertise. The differences are real: optimising to be quoted rather than clicked, entity clarity mattering more than keyword placement, corroboration across third-party sources carrying unusual weight, and measurement that cannot rely on a rank tracker. We sell them separately because the work is different, and we will tell you if you only need one.

Can you guarantee ChatGPT will cite my business?

No. These systems re-rank on every query, vary between users and sessions, and are opaque from the outside. Anyone offering a guarantee is selling something they cannot control. What we commit to is the foundation work, a tracked prompt panel so you can see movement, and honest reporting when a month shows none.

Does llms.txt actually do anything?

Not demonstrably, yet. It is a proposed convention rather than an adopted standard, and no major assistant has confirmed using it as an input. We ship it anyway because it takes about an hour, carries no downside and may become meaningful. If an agency is leading their pitch with llms.txt, that tells you how deep the rest of the work goes.

Should I block AI crawlers?

It depends on what your content is for. If your pages are the product — a paywalled archive, a subscription research library — blocking training crawlers is defensible. If your pages are marketing for a service you sell, blocking the crawlers that feed the systems your buyers now ask is working against yourself. The controls are also more granular than people realise: training crawlers and search crawlers are often separate agents and separate decisions. We lay out the options and you decide.

How long before this shows anything?

Technical and entity foundations can be in place in weeks, and fixing incorrect facts about your business sometimes resolves quickly. Becoming a consistently cited source takes months, because the corroboration layer is slow by nature. Expect leading indicators — crawler access, schema health, prompt-panel appearances — well before anything shows in revenue.

What does it cost, and does it replace my SEO budget?

It starts at $800 per month and it should sit alongside SEO rather than replace it. Organic search is still where most of the traffic is; this protects you against the share of it that is being absorbed into answers. If budget forces a choice and your technical foundation is weak, fix the foundation first — it is a prerequisite for both.

Find out what the assistants say about you

We will run a short prompt panel on your category and your brand across the major assistants and send you the raw results — who gets named, what is said about you, and whether any of it is accurate.