These definitions explain how this guide uses common AI search terms. Follow the linked chapters for methods, examples, and limitations.
A to Z
- AEO (answer engine optimization)
- Work intended to improve appearance in direct answers, including featured snippets and generated answers. Usage varies and overlaps with other optimization terms. See AI Search Terminology: SEO, GEO, AEO, and LLMO.
- Agent (agentic AI)
- An AI system that can use tools to perform tasks, such as browsing or completing a supported transaction. Capabilities and required user approvals depend on the system. See Products, Comparisons, and Buying Tasks, Changes to Watch in AI Search.
- AI Mode
- Google’s conversational search experience with generated responses and supporting links. Its behavior can differ from AI Overviews. See AI Search Products and Their Sources.
- AI Overview
- A generated summary shown for some Google searches, with links to supporting pages. See AI Search Products and Their Sources.
- AI search
- A search or information-seeking experience that uses an AI-generated response. It may also show conventional results and other interface elements. See What Is AI Search?.
- AI search optimization
- The work intended to improve how a business appears in AI search. This guide includes SEO and paid placement in that scope, while measuring paid and organic outcomes separately.
- AI search visibility
- How a brand or source appears in the tested AI answers, including mention frequency, recommendation frequency, prominence, and accuracy. See Measuring AI Search Visibility.
- CCBot
- Common Crawl’s crawler. Common Crawl publishes web datasets that can be used for research and model development. A fetched page is not guaranteed inclusion in any specific training dataset. See Crawlers, Indexes, and Training Data.
- ChatGPT-User
- An OpenAI user agent associated with actions made for users. Its purpose and access behavior are documented separately from GPTBot and OAI-SearchBot. See Crawlers, Indexes, and Training Data.
- Chunk
- A portion of a document used as a unit for processing or retrieval. Some systems use passages, others whole documents, or a combination. See Structuring Sites and Pages for Retrieval.
- Citation
- A reference identifying a source for an answer or claim. A citation does not necessarily expose every source considered during retrieval. See What Actually Correlates with AI Search Visibility.
- ClaudeBot, Claude-SearchBot, Claude-User
- Anthropic’s documented agents for different content-access uses. Check the current provider documentation before setting access rules. See Crawlers, Indexes, and Training Data.
- Context window
- The model’s limit on the input and generated content considered for a request, commonly measured in tokens. Exact accounting depends on the model.
- Crawler
- Software that requests web content automatically. A user-agent string can identify a claimed crawler but does not authenticate it. See Crawlers, Indexes, and Training Data.
- Embedding
- A numerical vector produced from an input. Comparing vectors can help estimate similarity for tasks such as retrieval. See Embeddings: How Machines Represent Meaning.
- GEO (generative engine optimization)
- A term introduced in a 2023 research paper for improving source visibility in generated answers. Broader industry usage varies. See AI Search Terminology: SEO, GEO, AEO, and LLMO.
- Google-Extended
- A Google robots.txt control token governing specified uses for Gemini training and grounding. Its documented scope is separate from Google Search inclusion and ranking. See Crawlers, Indexes, and Training Data.
- GPTBot
- OpenAI’s crawler associated with content that may be used for model training. Its control is separate from OAI-SearchBot. See Myths and Misconceptions.
- Grounding
- Supplying or connecting an answer to external information. Grounding can support verification, but does not guarantee accuracy or complete citations. See Retrieval and Grounding: How an AI Answer Gets Assembled.
- Hallucination
- A generated statement that is false or unsupported. Fluent wording does not establish that the statement is correct.
- Knowledge cutoff
- A provider-stated limit on a model’s training knowledge. It does not guarantee knowledge of every earlier fact or prevent use of newer information supplied in context. See Keeping Information Current.
- Knowledge graph
- A structured representation of entities and relationships between them. See Clear and Consistent Business Information.
- LLM (large language model)
- A model trained to process and generate language. Text generation commonly proceeds by selecting tokens in sequence. See How an LLM Turns a Prompt into a Response.
- LLMO (large language model optimization)
- A term for work intended to affect how language models describe or recommend a subject. Its scope varies by author. See AI Search Terminology: SEO, GEO, AEO, and LLMO.
- Mention rate
- The percentage of completed sampled answers containing the brand, under a stated matching rule. See Measuring AI Search Visibility.
- OAI-SearchBot
- OpenAI’s crawler associated with ChatGPT search. Allowing access does not guarantee inclusion or citation. See Crawlers, Indexes, and Training Data.
- Parametric memory (model memory)
- Information learned in model parameters during training. It is distinct from documents or other information supplied during a request. See Model Memory and Live Retrieval.
- Persona
- A description of a buyer or user included in a test prompt. It tests supplied requirements rather than reproducing all features of real account personalization. See Measuring AI Search Visibility, Competitive Analysis.
- PerplexityBot, Perplexity-User
- Perplexity agents associated with search collection and user-requested access. Consult current documentation for their exact purposes and rules. See Crawlers, Indexes, and Training Data.
- Prominence
- Where and how noticeably a brand appears in a response, measured under a defined rule. Keep position observations separate from mention frequency.
- Prompt
- Input asking an AI system to perform a task. For measurement, preserve the exact text and its version. See Prompt and Topic Research.
- Query fan-out
- Issuing several related search queries to address one user request. Available logs and result attribution vary by product. See Retrieval and Grounding: How an AI Answer Gets Assembled.
- RAG (retrieval-augmented generation)
- An approach that retrieves information and supplies it to a generative model for the response. See Retrieval and Grounding: How an AI Answer Gets Assembled.
- Ranking factors
- Signals used to order results. Research may also examine factors associated with generated recommendations. A correlation alone does not establish that a provider uses the factor causally. See What Actually Correlates with AI Search Visibility.
- Recommendation rate
- The percentage of completed sampled answers that recommend a brand under a stated scoring rule.
- Retrieval (live retrieval)
- Obtaining information from a source during a request or task. The source can be a search index, document collection, database, or another service. See Model Memory and Live Retrieval.
- robots.txt
- A site-root file specifying access rules for cooperating crawlers under the Robots Exclusion Protocol. It is not authentication or a method for deleting previously collected data. See Technical Accessibility for AI Crawlers.
- Share of voice
- A brand’s counted mentions divided by all counted brand mentions in the sample. Define deduplication and weighting. It is not market share.
- Temperature
- A generation setting that changes the distribution used for token selection where supported. It is one possible contributor to response variation.
- Token
- A unit used to represent model input or output. Text tokens may represent words, word parts, punctuation, or other sequences. Size depends on the tokenizer and language.
- Training data
- Examples used to adjust a model’s parameters. Inclusion and retention of a specific fact cannot be inferred simply from a crawler request. See Crawlers, Indexes, and Training Data.
- Volatility
- Variation in measured answers across stated conditions or periods. Separate repeated-run variation from deliberate prompt or model changes.