Short answer: AI search has brought a wave of new vocabulary, but most terms describe familiar ideas: how answers are produced (retrieval, grounding, citations), how content is accessed (crawlers, tokens, robots rules), how it is presented (AI Overviews, AI Mode, answer engines) and how visibility is measured (mentions, AI referrals, zero-click). This glossary defines 30 common terms in plain English, with notes on why each matters for a website.
How to use this glossary
The terms are grouped by topic rather than alphabetically, so related ideas sit together. Each definition is short and practical: what the term means and why a site owner might care. Where a term is a proposal or a marketing label rather than an established standard, the definition says so. Terms and products in this field change quickly, so treat operator documentation as the final word on specific product behaviour.
If you are new to the topic, start with the first two groups. They cover the ideas that affect what you should check on your own site.
Keep this page bookmarked as a reference when reading reports, proposals or articles about AI search. When a term appears without explanation, a quick check here helps you judge whether the writer means something specific.
Answers and interfaces
- AI search: any search experience where an AI system writes an answer based on retrieved sources, rather than only listing links.
- Answer engine: a product designed mainly to answer questions directly with sources, such as Perplexity. The term predates generative AI and was also used for voice assistants and featured snippets.
- AI Overviews: Google’s AI-generated summaries shown at the top of some search results pages, with links to supporting pages.
- AI Mode: a conversational mode in Google Search for complex questions, which runs several related searches and lets users ask follow-ups.
- Chat assistant: a general AI assistant, such as ChatGPT, Claude, Gemini or Copilot, that can answer from its training or search the web when needed.
- Citation: a link or reference to a source shown alongside an AI answer. Being cited is the AI search equivalent of appearing as a result, and it can bring both visibility and visits.
- Zero-click search: a search where the user gets the answer on the results page and does not visit any website.
Why these matter: the interface determines how your content is shown and whether a click follows. AI Overviews and AI Mode are part of Google Search and follow its rules, while chat assistants and answer engines have their own crawlers and sources. Knowing which surface a customer uses tells you which controls and indexes are relevant.
Crawling and access
- Crawler: an automated program that fetches web pages. Search engines and AI companies run their own crawlers, each with a user-agent name.
- User agent: the name a crawler or browser sends with each request, such as GPTBot or Googlebot. Robots rules and logs use it to identify visitors.
- Training crawler: a crawler that collects content for training AI models, such as GPTBot or CCBot.
- Search crawler (AI): a crawler that builds an index for an AI assistant’s search feature, such as OAI-SearchBot or PerplexityBot.
- User-triggered fetcher: an agent that fetches a specific page because a user asked, such as ChatGPT-User. Several operators say these may not follow robots.txt the same way as crawlers.
- robots.txt: a file at the root of a domain telling crawlers which paths they may fetch. It is a standard (RFC 9309), but compliance is voluntary.
- Product token: a robots.txt name that controls a use of content rather than a crawler, such as Google-Extended or Applebot-Extended.
- llms.txt: a proposed Markdown file listing a site’s key pages for language models. It is a community convention, not an official standard, with limited confirmed adoption.
Why these matter: access terms describe the part of AI search you control most directly. A single rule in robots.txt, or a CDN setting that blocks a user agent, can decide whether an assistant can use your pages at all. Most audits of AI visibility start here, and many problems are found here.
How the systems work
- Large language model (LLM): an AI model trained on large amounts of text to understand and generate language. It powers the writing part of AI answers.
- Training data: the text a model learned from. It has a cut-off date, so it can be outdated.
- Retrieval-augmented generation (RAG): a method where a system retrieves relevant documents and gives them to a model to write an answer. Most AI search uses some form of it.
- Grounding: tying an AI answer to specific sources, so statements are based on retrieved content rather than memory alone.
- Query fan-out: Google’s term for splitting one question into several related searches and combining the results.
- Embedding: a numerical representation of text meaning, used to find passages that match a question even when words differ.
- Chunk or passage: a section of a page that is retrieved and used on its own. Clear, self-contained paragraphs under descriptive headings make better chunks, because they still make sense when read in isolation.
- Hallucination: an AI answer that states something false with confidence. Grounding reduces it but does not eliminate it.
- AI agent: an assistant that can browse websites and take actions, such as filling in forms, on a user’s behalf.
Why these matter: understanding retrieval explains most practical advice. Because systems work with chunks, clear headings and self-contained paragraphs help. Because answers are grounded in retrieved sources, accurate and current pages matter. Because training data has a cut-off, fixing an error on the web may take time to reach every answer.
Optimisation and content
- GEO (generative engine optimization): making content easy for AI answer engines to find, understand and cite. It overlaps heavily with SEO.
- AEO (answer engine optimization): optimising content to be the direct answer, originally in snippets and voice results.
- Answer-first writing: placing a direct answer at the start of a page or section, then adding detail.
- Snippet controls: directives such as nosnippet, max-snippet and data-nosnippet that limit how text is shown, including in Google’s AI features.
Why these matter: these labels describe goals and techniques rather than standards. They are useful shorthand in conversation, but the work behind them is mostly familiar: clear structure, direct answers, specific facts and careful use of controls that limit how your text is shown.
Measurement
- AI referral traffic: visits that arrive from AI assistants, identified by referrer or source domains such as chatgpt.com or perplexity.ai. It is usually undercounted.
- Brand mention rate: the share of a fixed set of questions for which an assistant names your business. A practical substitute for rankings in AI answers.
Why these matter: without rankings, measurement relies on several imperfect signals. Referral data shows clicks, mention rates show presence in answers, and crawler logs show access. None is complete alone, which is why good reports combine them and explain their limits.
Terms to treat with caution
Some expressions circulate widely but promise more than they deliver. “AI ranking position” suggests a stable place in answers that does not exist. “AI visibility score” can be useful if the method is transparent, but it is misleading when built on a handful of prompts. “Special AI schema” does not exist as a requirement; standard structured data supports understanding but does not force inclusion. When you meet a new term, ask what concrete check or change it refers to. If nobody can say, it is probably a label rather than a method.
| Term | Status |
|---|---|
| robots.txt | Internet standard (RFC 9309) |
| Schema.org structured data | Widely supported vocabulary |
| llms.txt | Community proposal |
| GEO, AEO, LLMO | Industry labels, not standards |
| “AI ranking” | Misleading; answers have no stable positions |
How Site SEO AI Audit helps
Several terms in this glossary correspond to concrete checks in a Site SEO AI Audit report. Its AI visibility area checks AI crawlers in robots.txt, llms.txt, content that needs JavaScript, and the structure and dates AI answers quote, while other areas cover indexing, canonicals, structured data and speed. You can run a free audit to see these ideas applied to your own site.
Related reading
- How AI Search Engines Work: Retrieval, Ranking and Answers
- SEO vs GEO vs AEO: What the Terms Really Mean
- How to Allow or Block AI Crawlers in robots.txt
The bottom line
Most AI search vocabulary describes ideas site owners already know: crawling, indexing, answering and measuring. Knowing the precise meaning of terms such as training crawler, product token, grounding, query fan-out and citation helps you ask better questions, spot empty promises and focus on the checks that actually change what AI systems can see on your site.
DUK
What is the difference between a training crawler and a search crawler?
A training crawler collects content to train AI models, while a search crawler builds an index that an assistant searches when answering. Blocking one does not necessarily block the other.
What does grounding mean in AI search?
Grounding means basing an AI answer on specific retrieved sources rather than only on the model’s memory. It makes answers more current and allows citations.
Is llms.txt a standard?
No. It is a community proposal for a Markdown file that lists a site’s key pages for language models. Support among major AI search products is limited and unconfirmed.
What is query fan-out?
It is Google’s term for breaking one question into several related searches, used in AI Mode and AI Overviews, and combining the results into one answer.
What is a zero-click search?
It is a search where the user finds the answer on the results page and does not click through to any website, for example from an AI Overview or a snippet.


