Matt Hardy

GUIDE

Answer engine optimisation

Search used to end with a list of links you could count your position in. It increasingly ends with a paragraph that either mentions you or doesn’t. This is how I measure that, and what twelve months of doing it changed my mind about.

Matt Hardy · Growth marketing · Auckland, New Zealand

01

What it is, and what it is not

Answer engine optimisation is the work of being cited accurately in the generated answers that ChatGPT, Perplexity, Google’s AI Overviews, Claude and Copilot return instead of a list of links. The unit of success is a citation, not a position.

It is not a replacement for search engine optimisation and anyone selling it as one is selling something. Every answer engine that cites live sources is retrieving those sources from an index — usually a search index. If a page cannot be crawled, cannot be parsed, or does not exist, no amount of AEO technique will get it quoted. Classic technical SEO is the floor. AEO is what you do standing on it.

The difference that matters is what the machine does with the page once it has it. A search engine ranks the document. An answer engine extracts a claim from the document, reattributes it, and drops the rest. So the question stops being “does this page rank for the query” and becomes “is there a sentence on this page that a model can lift, attribute, and be right about”.

The whole discipline follows from that one shift: from ranking documents to sourcing sentences.

02

The measurement problem

There is no rank. That is the first thing to sit with, because every reporting habit you have is built on there being one.

Answers are generated, so they vary by phrasing, by session, by model version, and often by nothing you can see. Ask the same question twice and you can be cited once. There is no position 4 to hold and no ranking API to query. If you report on answer engines the way you report on Google, you will produce a number that moves for reasons you cannot name and you will be asked to explain it.

So you sample. You define a prompt set, run it on a schedule, capture the answers and the citations, and treat the result as a survey rather than a scoreboard. A single run tells you almost nothing. A hundred runs a week for three months tells you where you actually stand, and — more usefully — who is standing there instead of you.

03

Share of voice, not position

The metric that replaces rank is share of voice: across your prompt set, in what proportion of generated answers is your domain cited at all, and what proportion of the total citations in those answers are yours.

It behaves like a survey statistic because that is what it is, which means it comes with a margin of error and needs a sample size before it means anything. That is unfamiliar and it is also the honest version. A share-of-voice figure with no prompt count next to it is decoration.

Two refinements earn their complexity. Weight prompts by how commercially real they are, because being cited on a definitional query and being cited on a comparison query are not the same event. And record the competitor set in every answer, because the useful finding is almost never your own number — it is that the same three domains are answering your category and one of them is a forum post from 2019.

Citation rate
Share of sampled answers that cite the domain at all.
Citation share
Share of all citations across those answers that are yours.
Weighted share
The same, with commercial-intent prompts counted heavier than definitional ones.
Competitor set
Every other domain cited, recorded per answer. The most useful column in the table.

04

The five systems

This is the shape of the pipeline described in panel 03 — five systems, twelve months of data. Nothing here needs a vendor. It needs a prompt set, somewhere to put rows, and the discipline to run it on a schedule after the novelty has worn off.

  1. The prompt setFifty to two hundred real questions, versioned like code. Buyer language, not keyword language. Frozen between measurement periods or the trend line is meaningless.
  2. CaptureRun the set across each engine on a schedule. Store the full answer text, every cited URL, the model, and the timestamp. Store the answer, not a verdict about it — you will want to re-score old runs against new questions.
  3. Share of voiceRoll capture up into citation rate, citation share and the competitor set, per engine and per prompt cluster. This is the number that goes in front of anyone senior.
  4. Gap rankingFor every prompt where you are not cited, what is cited instead, and could you plausibly displace it. Ranked by winnability against effort, not by volume. Most gaps are not worth closing and the ranking is what says so.
  5. The fix queueThe ranked gaps as actual work: a page to write, a claim to make quotable, a schema block to add, a crawler to unblock. If the pipeline does not end in a queue somebody works through, it is a dashboard.

Four of these are plumbing. The fourth is the one with judgement in it, and the one worth your time.

05

What actually moves citations

Ranked roughly by how reliably it worked, which is not the order the discipline usually presents them in.

  1. Being crawlable by the retrievers specifically.Answer-engine crawlers split into two roles that get conflated: trainers and indexers such as GPTBot, ClaudeBot, PerplexityBot and Google-Extended, and live retrievers such as ChatGPT-User, Claude-User, OAI-SearchBot and Perplexity-User. The retrievers are what fetches your page at the moment somebody asks a question about it. Plenty of sites block all of them by default, discover it a year later, and find that the entire programme was a robots.txt line.
  2. Making one claim per passage, stated plainly, near a heading that matches the question.Extraction is the mechanism, so write the sentence you would want quoted, and put it where a model reading the section will find it first. This is not a trick — it is the same edit that makes the page better for a human skimming it.
  3. Being the primary source of a number.A specific, attributable figure that exists nowhere else gets cited, and keeps getting cited, long after the page that carries it stops ranking. Original data outperforms restated data by a distance.
  4. Structured data and clean semantics.Worth doing, cheap to do, and consistently less decisive than the three above. A page that cannot be crawled does not benefit from perfect schema.
  5. Being mentioned elsewhere.Answer engines retrieve from indexes that reward corroboration, so the pages that get cited tend to be pages other people already reference. The old work still counts.

06

What the data changed my mind about

I expected the winnable gaps to be the high-volume ones. They were not. The prompts where displacement actually happened were narrow, specific, often low-volume questions where the incumbent citation was weak — a thin listicle, an unmaintained forum thread, a competitor’s outdated help doc. The ranked list disagreed with my instincts. I shipped it anyway, and that was the right call.

I also expected volatility to settle. It did not. Month-to-month variance stayed high enough that any single reading was noise, and the only defensible way to report was a rolling window with the sample size printed next to it. Anyone promising you a stable AEO rank is not measuring one.

And the least glamorous finding: the largest single improvement in the whole twelve months came from unblocking crawlers, not from writing anything. Check that first. It costs an afternoon and it is embarrassingly often the answer.

If you take one thing from this page: sample properly, rank the gaps honestly, and read your robots.txt before you write a word.

07

Common questions

What is answer engine optimisation?
Answer engine optimisation is the practice of getting a site cited accurately inside the generated answers produced by AI systems such as ChatGPT, Perplexity, Google’s AI Overviews and Copilot. Unlike search engine optimisation, where the unit of success is a ranking position, the unit of success in AEO is a citation in a generated answer.
Is AEO different from SEO?
Yes, but it depends on it. Answer engines retrieve their sources from search indexes, so a page that cannot be crawled or parsed cannot be cited no matter how it is written. Technical SEO is the prerequisite. AEO is the additional work of making individual claims on the page extractable and attributable once the page is retrievable.
How do you measure AEO?
By sampling rather than ranking. Define a fixed prompt set of real buyer questions, run it across each answer engine on a schedule, and record every citation returned. The resulting metric is share of voice: how often your domain is cited across sampled answers, and what proportion of total citations are yours. A single run is noise; a rolling window over months is a measurement.
Do I need to allow AI crawlers in robots.txt?
To be cited by live retrieval, yes. Crawlers such as ChatGPT-User, Claude-User, OAI-SearchBot and Perplexity-User fetch pages at the moment a user asks a question, and blocking them makes the page uncitable in that moment. Training crawlers such as GPTBot, ClaudeBot and Google-Extended are a separate decision with different trade-offs, and can be allowed or blocked independently.
How long does AEO take to show results?
Long enough that you need the measurement running before the work starts. Answer output varies enough between runs that a change is only visible against a baseline of several months of sampling, so the first useful deliverable is the pipeline itself rather than any content change.

Built one of these?

I have run this pipeline for twelve months and would rather compare notes than pitch anyone. If you are measuring answer engines and getting numbers you cannot explain, the explanation is usually sampling.

qmmhardy@gmail.comBack to the portfolio