Beta pricing ends when monitoring launches. Current customers keep their price for 12 months. Lock it in →

Glossary
RAGRetrieval Augmented Generation

What is Retrieval-Augmented Generation?

Retrieval-Augmented Generation (RAG) is the technique of fetching relevant documents at the moment a question is asked and having the model write its answer from those documents. It is how every major AI answer engine works, which is why AI visibility is mostly a retrieval problem.

Retrieval-Augmented Generation is a two-step arrangement. When a question arrives, the system first searches a body of documents for passages that look relevant, then hands those passages to a language model and asks it to write an answer using them. The term comes from a 2020 paper by Lewis and colleagues at Facebook AI Research, presented at NeurIPS, which combined what it called parametric memory, the knowledge baked into the model's weights, with non-parametric memory, an external collection the model can look things up in.

The plain version: a model without retrieval is sitting a closed-book exam from memory. A model with retrieval is sitting an open-book exam, and the book is whatever the retrieval step handed it.

For marketers, one fact does most of the work here. Google's AI Overviews and AI Mode, ChatGPT with browsing, Perplexity, Gemini with grounding and Claude with web access are all variants of this arrangement. They are not answering from memory alone. They are fetching pages at question time and writing from what they fetched.

That reframes the most common question people ask about AI visibility. "How do we get into the training data" is largely the wrong question, and it is not one you can act on anyway: training runs are periodic, closed, and cannot be edited retroactively. Retrieval happens fresh on every query, which means a page published this week can be cited this week. It also means the gate is the retrieval step, not the writing step. If your page is not retrieved, nothing about how well it is written matters, because the model never sees it. Crawlability, being fetchable by the right crawlers, and ranking well enough to be in the candidate pool are floor conditions rather than refinements.

Two honest caveats. First, training data has not become irrelevant: when someone asks a model for recommendations with no retrieval in play, it answers from what it absorbed, and that shapes baseline brand recall. Both layers exist and the sensible strategy feeds both. Second, RAG is a family of designs rather than one fixed pipeline. Systems differ in what they search, how many candidates they pull, how they rank them and how much of each document reaches the model, which is a large part of why the same query produces such different source lists on different engines.

You do not need to understand the machinery to act on it. You need to know that being retrievable comes first, and that it is a separate problem from being persuasive.

See where your site stands on this today.

Run a free website audit to see exactly how your site scores, with evidence, not just a definition.