Insights · mechanism guide

How answer engines choose which sources to cite.

Answer engines choose their sources through three filters applied in sequence: retrieval (can the engine find and read you?), trust-weighing (does the evidence let it rely on you?) and synthesis (are you worth naming in the final answer?). A business can fail at any filter, and each failure has a different fix. Understanding the mechanism is what separates deliberate optimisation from guesswork.

Filter one: retrieval - can the engine reach and read you?

Before anything can be cited it must be found. When an engine handles a question it retrieves candidate sources: from its training knowledge, from live web search, or from an index built by its own crawlers - GPTBot, ClaudeBot, PerplexityBot and the rest. Two failures are common here, and both are invisible from the inside.

The first is the silent block. Many websites block AI crawlers at the firewall or CDN edge - often as a security default nobody chose deliberately. A site can explicitly allow AI crawlers in its robots.txt and still return an error to them at the edge, because the two controls are independent. The business believes it is open; the engines have never read a page of it.

AI crawler CDN edge robots.txt Rendered HTML Structured data Knowledge graph ↑ often silently blocked here
fig. 1 - two independent gates: a site can pass robots.txt and still be blocked at the edge.

The second is unreadability. A page the crawler reaches can still be opaque: content locked inside scripts, meaning conveyed only visually, no structured data to declare what anything is. The engine retrieves it and learns little.

Filter two: trust - does the evidence support relying on you?

Retrieval produces candidates; trust decides which survive. An engine putting a recommendation in front of a user is taking a small reputational risk, and it manages that risk the way a careful human would: by favouring sources it can verify and cross-check.

Three kinds of evidence carry the weight. Consistency - the business presents as one coherent entity, with the same name, location and facts on its website, its Google Business Profile, review platforms, the company register and LinkedIn. Corroboration - independent third parties mention, list and cite the business without being paid to; engines lean hard on sources with no stake in the outcome. Credentials - the content is attached to named, verifiable people, which matters most in health, legal and financial topics where engines are visibly more conservative.

Filter three: synthesis - are you worth naming?

Finally the engine writes its answer, and here the practical question is quotability. Engines favour sources that make accurate lifting easy: a direct answer in the opening lines, specific facts with real numbers, clean question-and-answer structure, honest freshness dates. Between two equally trustworthy sources, the one that states its answer plainly gets cited; the one that buries it in marketing prose gets paraphrased away - or skipped.

Why do answers and recommendations behave differently?

Because they lean on different filters. A direct answer (“what does a dental implant involve?”) is mostly a synthesis problem: the engine needs a clear, liftable explanation, so machine-readability dominates and structured data punches above its weight. A recommendation (“who is the best implant dentist near me?”) is mostly a trust problem: the engine is staking its judgment, so corroboration and credentials dominate. In our own local-market testing this split shows up consistently: schema gets a business cited; independent authority gets it recommended. Optimising only one side leaves the other half of your visibility on the table.

Where is your business failing?

The three-filter diagnostic

  1. Absent everywhere, even on brand questions? Suspect retrieval. Check AI-crawler access at the edge, robots.txt, and whether key content is readable without scripts.
  2. Retrieved but never recommended - or described inaccurately? Suspect trust. Check entity consistency across every profile, the depth of independent citations, and whether real people with credentials stand behind the content.
  3. Named for some questions but losing the ones that matter commercially? Suspect synthesis on those pages - and remember visibility is won question by question. Make the losing pages answer-first, factual and quotable.

Each filter is testable from the outside, which is why an evidence-first audit can locate the failure precisely before any money is spent fixing the wrong thing. The full repair sequence is the Answari methodology.

Stacey Smith

Stacey Smith

Founder, Answari

Ten years in brand, search and performance marketing. Stacey founded Answari to help UK professional practices get found, chosen and cited by AI answer engines. About Stacey →