Skip to main content

How to Learn RAG (Retrieval-Augmented Generation)

Retrieval-augmented generation is a fine-tuning recipe that lets a generator condition on text it has just retrieved. Patrick Lewis and colleagues define it in Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks as models that combine parametric memory, a pre-trained sequence-to-sequence network, with non-parametric memory, a dense vector index of Wikipedia read by a neural retriever. The weights hold what the generator already knows. The index holds passages you can open and replace. A generator that answers from weights alone can still hallucinate, and provenance for that answer stays hard to show. Their remedy is to retrieve passages and generate from them.

Learn it in the order that paper builds the system: the retrieve step, the documents, the generator, then what breaks.

The retrieve step

The input is a sequence. The retriever scores passages and keeps the top K. Lewis et al. use Dense Passage Retrieval: one BERT-base encoder for the query, one for each document, and the score is the inner product of the two vectors. The lookup is approximate maximum inner product search.

They initialize from a DPR model trained to fetch passages that contain answers to TriviaQA and Natural Questions. Training does not label a correct passage. The retrieved document is a latent variable, and the loss is the negative log probability of the target after marginalizing over the top passages. They fine-tune the query encoder and the generator. The document encoder stays frozen, because updating it means rebuilding the index.

RAG-Sequence uses one retrieved document for every output token. RAG-Token may use a different document for each token. Jeopardy questions in the paper often pack two facts, and RAG-Token, which can draw them from two passages, outperforms RAG-Sequence there.

The documents

The experiments use the December 2018 Wikipedia dump. Each article is split into disjoint 100-word chunks, 21 million documents. Those chunks are embedded and stored in a FAISS index with a Hierarchical Navigable Small World graph. The 100-word split is how this paper prepared Wikipedia. It is the setting they measured.

Each chunk stays readable text. The index stores its vector, and the index can be replaced. They build a second index from the December 2016 dump and ask "Who is {position}?" about 82 world leaders who changed between the dumps. The 2016 index answers 70% of the 2016 leaders correctly. The 2018 index answers 68% of the 2018 leaders correctly. The 2018 index on the 2016 leaders scores 12%. The 2016 index on the 2018 leaders scores 4%. Swapping the index updates the answers without retraining the generator.

The generator

The generator is the parametric memory. Any encoder-decoder can fill the role. They use BART-large and concatenate the input with the retrieved passage. The next token depends on that pair and on the tokens already written.

In the Hemingway example, the document posterior is high for the passage that mentions "The Sun Also Rises" while that title is generated, then flattens. BART alone finishes the title from its weights. The passage steers. The parameters complete a name they already store. A passage that never contains the answer verbatim can still move the output, which is how generation in the paper beats an extractive reader that needs the span in the text. On Natural Questions, RAG still scores 11.8% accuracy when the answer is in none of the retrieved documents. An extractive model would score 0% there.

What breaks

Retrieval can collapse. In preliminary story-generation runs the retriever returned the same documents for every input. The generator learned to ignore them, and RAG matched BART. Lewis et al. tie that to tasks with a weaker need for a specific fact, and to long targets that give the retriever a weak gradient.

A frozen retriever is weaker. Learning to retrieve improves results on all of their tasks. Dense retrieval still loses on FEVER, where a fixed BM25 word-overlap retriever scored higher, because the claims are full of entity names. Dense retrieval is what helps the other tasks, especially open-domain question answering.

More documents at test time help up to a point. RAG-Sequence's Natural Questions score rose as more passages were retrieved. RAG-Token peaked at 10. Training with 5 latent documents or with 10 did not show a significant difference.

The corpus limits the answers. The broader-impact section says Wikipedia, or any external source, will probably never be entirely factual or free of bias. Some MS MARCO questions, such as "What is the weather in Volcano, CA?", cannot be answered from that dump, and the model falls back on its weights. An empty "null document," for inputs where nothing useful was retrieved, did not improve their results, so they left it out.

What to practice

For one question, mark whether the answer is written in a retrieved passage, only implied, or missing. That is the split behind the 11.8% figure.

Ask the same question against a second index. If the answer stays put, you are reading the generator. If every query returns the same passages, that is retrieval collapse.

On claim-checking questions, compare a word-overlap retriever with the dense one. FEVER is where overlap won. Open-domain questions are where the dense retriever mattered.

A course from the paper

Ailurn turns a prompt, a PDF, a GitHub repo, or a docs URL into a course you take in the same workspace. It does not watch YouTube videos, and it does not issue accredited certificates. A PDF of the Lewis et al. paper, or that paper's URL, can ground the course. Start from the AI course builder and ask for a course on retrieval-augmented generation: the DPR retrieve step, a document index you can swap, RAG-Sequence versus RAG-Token, and a check for retrieval collapse and a stale index.

Start

Name a subject. Leave with a course.

Design, finance, math, interviews, or code. You create it. You learn it here.