Retrieval-Augmented Generation (RAG)

If you are learning Generative AI, one term you will hear again and again is RAG — Retrieval-Augmented Generation.

At first, the name sounds complicated. But the basic idea is surprisingly simple.

RAG solves one fundamental problem ↓

Prefer watching the video version?

Video thumbnail for the Day 12 lesson: Retrieval-Augmented Generation (RAG)
Retrieval-Augmented Generation (RAG) — Video Tutorial Watch on YouTube ↗
How can an LLM answer questions using information that was not part of its original training data?

For a DevOps engineer, this is especially important.

Your company may have:

Kubernetes runbooks
Terraform documentation
Helm charts
Internal wikis
Incident reports
Postmortems
AWS architecture documents
GitHub repositories
Troubleshooting guides

An LLM does not automatically know all of this information.

RAG gives us a way to retrieve relevant information from our own knowledge base and provide it to the LLM before it generates an answer.

RAG overview: encode documents into a vector database, encode the query, run similarity search, combine similar documents with the query, and generate a final response with an LLM.
Documents are indexed offline. At query time we retrieve similar chunks, add them to the prompt, and ask the LLM to answer.
The problem

1. Why do we need RAG?

Imagine you ask an LLM:

Why is our payment-api Pod restarting continuously?

The model might know Kubernetes. It probably knows what a Pod is. It may know common reasons for CrashLoopBackOff.

But it does not automatically know your environment.

It may not know:

Your Kubernetes configuration
Your internal runbooks
Yesterday's production incident
Your company's architecture
Your Terraform modules
Your internal troubleshooting procedures

Without access to that information, the model can only answer based on the information already available to it.

Conceptually:

User
 │
 │ "Why is payment-api restarting?"
 ▼
LLM
 │
 ▼
General Kubernetes answer

The answer might be useful. But it is not necessarily based on your company's actual environment.

The idea

2. What if we give the LLM our own information?

Suppose your company has an internal runbook containing this:

payment-api may enter CrashLoopBackOff when it cannot
connect to the PostgreSQL database.

Check the DATABASE_HOST environment variable and verify
that the PostgreSQL Service is reachable.

Now imagine that before asking the LLM to answer, our application searches the company's documentation and finds this paragraph.

Then we send both:

User Question:
Why is payment-api restarting continuously?

Relevant Company Documentation:
payment-api may enter CrashLoopBackOff when it cannot
connect to PostgreSQL...

to the LLM.

Now the model has much better information. It can respond:

According to the internal runbook, payment-api may restart when it cannot connect to PostgreSQL. I would first verify DATABASE_HOST and check whether the PostgreSQL Service is reachable.

Notice what happened.

We did not retrain the LLM.

We simply retrieved useful information and gave that information to the model.

That is the core idea behind RAG.

Two words

3. What does RAG mean?

RAG stands for Retrieval-Augmented Generation.

Retrieval

Find information relevant to the user's question.

User:

"Why is my Pod restarting continuously?"

        ↓

Search company knowledge

        ↓

Find:

CrashLoopBackOff troubleshooting guide

Augmented Generation

Take the information we retrieved and add it to the information given to the LLM. Then ask the LLM to generate an answer using that information.

So:

Retrieve
   +
Generate
   =
RAG

Or, more accurately:

User Question
      │
      ▼
Retrieve Relevant Information
      │
      ▼
Add Information to Prompt
      │
      ▼
LLM Generates Answer

That is RAG.

Remember this

4. RAG in one sentence

RAG searches your knowledge base for information relevant to the user's question and gives that information to the LLM before the LLM generates its answer.
DevOps example

5. A DevOps example

Imagine your organization has:

Kubernetes documentation
Runbooks
Postmortems
Incident reports
Terraform documentation
AWS architecture documents
GitHub documentation

A user asks:

Why is Pod abc failing?

A RAG application could search:

Pod runbooks
Incident history
Known issues
Troubleshooting documentation

Then retrieve the most relevant information. Then send that information to the LLM.

User
 │
 │ "Why is pod abc failing?"
 ▼
Retriever
 │
 ├── Search runbooks
 ├── Search incidents
 └── Search known issues
 │
 ▼
Relevant Information
 │
 ▼
LLM
 │
 ▼
Answer

Now the answer can be grounded in your organization's knowledge instead of relying only on general LLM knowledge.

Critical concept

6. RAG has two major flows

This is one of the most important concepts to understand.

A RAG system usually has two very different flows:

Flow 1 — Ingestion

Prepare your documents so they can be searched efficiently.

Documents
    ↓
Load
    ↓
Split
    ↓
Create Embeddings
    ↓
Store in Vector Database
Data ingestion pipeline: Load documents, Split into chunks, Embed into vectors, Store in a vector database index.
Ingestion prepares your knowledge base so questions can be answered later.

Flow 2 — Retrieval and Generation

When the user asks a question:

Question
    ↓
Create Query Embedding
    ↓
Search Vector Database
    ↓
Retrieve Relevant Chunks
    ↓
Send Question + Chunks to LLM
    ↓
Generate Answer

Beginners often mix these two flows together. Keep them separate in your mind.

Ingestion happens when preparing or updating the knowledge base.

Retrieval happens when answering a question.

Offline pipeline

7. Let's start with the ingestion pipeline

Suppose your company has a 500-page Kubernetes troubleshooting manual.

You want your RAG application to answer questions using this manual.

We cannot simply throw a 500-page PDF into every LLM request.

Instead, we prepare the document.

PDF
 │
 ▼
Document Loader
 │
 ▼
Text
 │
 ▼
Chunking
 │
 ▼
Chunks
 │
 ▼
Embedding Model
 │
 ▼
Vectors
 │
 ▼
Vector Database

Let's understand every component.

Step 1

8. Document loader

The first job is simply:

Read the data.

Your data can come from many places:

PDF
TXT
Markdown
CSV
Web pages
Source code
Directories

A document loader reads information from these sources and converts it into something our application can process.

For example, if we have kubernetes-runbook.pdf, a PDF loader extracts the text:

Chapter 7: CrashLoopBackOff

CrashLoopBackOff occurs when a container repeatedly
starts and crashes...

Now our application has text that it can process.

Loaders

9. Different types of document loaders

Depending on the data source, different loaders can be used.

TextLoader — useful for .txt, .md, and source code.

PyPDFLoader — useful for PDF documents. It can extract text page by page.

CSVLoader — useful for structured CSV data.

DirectoryLoader — useful when you want to load many files from a directory.

WebBaseLoader — useful when information needs to be loaded from web pages.

Document loaders hub: TextLoader, PyPDFLoader, WebBaseLoader, DirectoryLoader, and CSVLoader.
The goal is the same: convert different data sources into documents the RAG pipeline can process.
Data quality

10. What about complicated PDFs?

Not every PDF is easy to process. Some PDFs contain:

Scanned pages
Screenshots
Tables
Diagrams
Forms
Multiple columns
Complex layouts

A simple PDF text extractor may not understand all of these correctly.

For more complicated documents, advanced document-processing tools can be useful because they understand document layout and can use OCR when necessary.

If document ingestion is bad:

Bad extraction
      ↓
Bad chunks
      ↓
Bad embeddings
      ↓
Bad retrieval
      ↓
Bad LLM answer

RAG quality begins with data quality.

Scale

11. Lazy loading

Imagine your organization has 10 files. Loading everything into memory is probably fine.

But what if it has 1,000,000 documents? Loading everything simultaneously may consume huge amounts of memory.

One approach is lazy loading.

Instead of loading every document and keeping everything in memory, we can process one document at a time:

Document 1
   ↓
Process

Document 2
   ↓
Process

Document 3
   ↓
Process

This can significantly reduce memory usage for large ingestion jobs.

Step 2

12. Chunking

Now we have extracted the text. But imagine our Kubernetes documentation contains 500 pages.

We normally do not want to treat those 500 pages as one giant piece of text.

Instead, we split the document into smaller pieces called chunks.

Large Kubernetes Document
          │
          ▼
        Split
          │
   ┌──────┼──────┐
   ▼      ▼      ▼
Chunk 1 Chunk 2 Chunk 3 ...
Chunk 1:
Kubernetes Pod Lifecycle

Chunk 2:
CrashLoopBackOff

Chunk 3:
ImagePullBackOff

Chunk 4:
OOMKilled
Why split?

13. Why do we need chunking?

First, embedding models have limits on how much text they can process.

Second, smaller chunks make retrieval more focused.

Suppose the user asks:

Why is my Pod restarting continuously?

The relevant information might occupy only one or two paragraphs in a 500-page document.

We don't want to retrieve all 500 pages. We want something closer to:

Question
   ↓
Find relevant chunks
   ↓
CrashLoopBackOff chunk
Trade-off

14. Chunk size matters

Imagine we make chunks extremely small:

Chunk 1: CrashLoopBackOff
Chunk 2: occurs when
Chunk 3: a container
Chunk 4: repeatedly crashes

We have destroyed the useful context.

Now imagine the opposite: one chunk with the entire 500-page manual. That is not useful either.

We want chunks that are:

Small enough to retrieve precisely

but

Large enough to preserve useful meaning
Boundaries

15. Chunk overlap

Sometimes important information sits exactly between two chunks.

One common technique is chunk overlap.

Instead of completely separate chunks:

Chunk 1: A B C D
Chunk 2: E F G H

we keep some overlapping content:

Chunk 1: A B C D
Chunk 2: C D E F
Chunk 3: E F G H

This can help preserve context across chunk boundaries.

Strategies

16. Different chunking strategies

Fixed-size chunking — split after a certain number of characters or tokens. Easy, but it does not necessarily understand document structure.

Recursive chunking — try to preserve natural boundaries such as sections, paragraphs, and sentences before splitting into smaller units. Often a practical starting point.

Semantic chunking — instead of splitting purely by size, split based on meaning so related information stays together.

LLM-based chunking — an LLM analyzes a document and decides where meaningful chunk boundaries should occur. More intelligent, but more complex and costly.

Five chunking strategies: fixed-size with overlap, semantic, recursive, document structure-based, and LLM-based.
Different strategies trade simplicity, structure awareness, and cost.
Ready to search

17. Now we have our chunks

After chunking, our Kubernetes documentation might look like:

Chunk 1 — Kubernetes Pod Lifecycle
Chunk 2 — CrashLoopBackOff troubleshooting
Chunk 3 — ImagePullBackOff troubleshooting
Chunk 4 — OOMKilled troubleshooting
Chunk 5 — Pod networking

Now we need a way to search these chunks based on their meaning.

This brings us to embeddings.

Step 3

18. Embeddings

The word embedding scares many beginners. The basic idea is actually simple.

An embedding converts text into numbers.

For example, "CrashLoopBackOff" might become something conceptually like:

[0.23, 0.11, 0.92, ...]

This list of numbers is called a vector.

The exact numbers are not important for understanding RAG.

The important part is what those numbers represent: they capture characteristics of the text's meaning in a numerical form.

Semantic search

19. Why convert text into numbers?

Suppose our document contains:

CrashLoopBackOff occurs when a container
repeatedly starts and crashes.

But the user asks:

Why is my Pod restarting continuously?

The user did not use the word CrashLoopBackOff.

A simple keyword search may struggle if it relies heavily on exact wording.

But semantically, these ideas are closely related. An embedding model converts both into vectors that can be close to each other because the meanings are similar.

Meaning space

20. Similar meaning → nearby vectors

Think of embeddings as placing meanings into a giant mathematical space.

              Kubernetes Problems

CrashLoopBackOff  ●
                  ● Pod keeps restarting


                                      ● Docker image layers


       ● PostgreSQL connection problem

Items with similar meaning tend to be closer. Items with unrelated meanings tend to be farther apart.

We can search based on meaning, not just exact words.
During ingestion

21. Embedding the documents

During ingestion, each chunk is sent to an embedding model.

Chunk 1 → Embedding Model → Vector 1
Chunk 2 → Embedding Model → Vector 2
Chunk 3 → Embedding Model → Vector 3

Now every document chunk has a numerical representation.

Storage

22. Where do we store all these vectors?

If you have 20 chunks, storage is easy.

But imagine a company has 100,000 PDFs, millions of documentation pages, thousands of GitHub repositories, years of incident reports, and thousands of runbooks.

After chunking, we may have millions of chunks. Each chunk has an embedding vector.

We need a system capable of efficiently storing and searching those vectors.

This is where the vector database comes in.

Vector database

23. What is a vector database?

A vector database is designed to store and efficiently search vectors.

Vector Database

Chunk 1 → [0.23, 0.11, 0.92, ...]
Chunk 2 → [0.82, 0.44, 0.05, ...]
Chunk 3 → [0.18, 0.15, 0.91, ...]

Along with the vector, we normally want to keep information about the original chunk:

Vector
Original text
Document ID
Source
Page
Version
Metadata

This allows us to retrieve the actual text after finding a matching vector.

Different search

24. Why not just use SQL?

Traditional databases are excellent at exact lookups:

SELECT *
FROM incidents
WHERE pod_name = 'nginx';

But RAG often needs a different type of search.

Instead of asking "Find this exact value," we want: "Find document chunks whose meaning is most similar to this question."

That is the type of search vectors are designed to help with.

Offline complete

25. Our ingestion pipeline is now complete

Company Documents
       │
       ▼
Document Loader
       │
       ▼
Extracted Text
       │
       ▼
Chunking
       │
       ▼
Document Chunks
       │
       ▼
Embedding Model
       │
       ▼
Vectors
       │
       ▼
Vector Database

This happens before the user asks the question. Now let's see what happens at query time.

Online starts

26. The user asks a question

The user asks:

Why is my Pod restarting continuously?

We cannot directly search the vector database using ordinary text. The database contains vectors.

Therefore, we first convert the user's question into a vector using an embedding model.

User Question
"Why is my Pod restarting continuously?"
              ↓
       Embedding Model
              ↓
[0.21, 0.13, 0.89, ...]

This is usually called the query embedding or query vector.

Top-K

28. Top-K retrieval

The vector database might find hundreds of possible matches. Should we send all of them to the LLM?

Usually, no. Sending all 200 matching chunks means more tokens, higher cost, longer response time, limited context window, more irrelevant information, and potentially worse answers.

Instead, we retrieve only the best few results. This is called Top-K Retrieval.

Choosing K

29. What does K mean?

K simply means: how many results should we retrieve?

Top-3 → best 3 chunks
Top-5 → best 5 chunks
Top-10 → best 10 chunks

Suppose K = 5. The remaining matches are ignored for this request.

Trade-off

30. Why not always use a large K?

More information is not always better.

With K = 50, many chunks may only be loosely related — higher cost, slower response, more noise, more conflicting information, reduced answer quality.

With K = 1, you may provide too little information if the guide is split across error description, root cause, and resolution chunks.

Choosing K is a trade-off. There is no universal value that is correct for every RAG application.

Prompt

31. Now we have the relevant context

Suppose Top-K retrieval returns:

Chunk 1: CrashLoopBackOff occurs when the application repeatedly exits.
Chunk 2: For payment-api, verify PostgreSQL connectivity.
Chunk 3: Check DATABASE_HOST and DATABASE_PORT environment variables.

The application can construct a prompt conceptually like:

Use the following company documentation to answer the question.

Context:
1. CrashLoopBackOff occurs when the application repeatedly exits.
2. For payment-api, verify PostgreSQL connectivity.
3. Check DATABASE_HOST and DATABASE_PORT.

Question:
Why is payment-api restarting continuously?

This entire prompt goes to the LLM.

Generate

32. The LLM generates the answer

The LLM now has its general Kubernetes knowledge + the user's question + relevant company documentation.

It can generate something like:

The internal documentation indicates that payment-api can enter CrashLoopBackOff when PostgreSQL connectivity fails. Start by checking DATABASE_HOST and DATABASE_PORT, then verify that the PostgreSQL Service is reachable from the Pod.

And again: we did not retrain the LLM. We retrieved relevant knowledge at query time and provided it as context.

Put it together

33. The complete RAG query flow

User
 │
 │ "Why is my Pod restarting?"
 ▼
RAG Application
 │
 ▼
Embedding Model
 │
 ▼
Query Vector
 │
 ▼
Vector Database
 │
 ▼
Similarity Search
 │
 ▼
Top-K Relevant Chunks
 │
 ▼
Prompt Builder
 │
 │ Question + Retrieved Context
 ▼
LLM
 │
 ▼
Answer
 │
 ▼
User

That is the basic RAG architecture.

Memorize this

34. Offline vs online flow

Offline / Ingestion

Documents → Load → Chunk → Embed → Store

Online / Query

Question → Embed → Search → Retrieve → Augment Prompt → LLM → Answer

A very short version:

INGESTION
Load → Split → Embed → Store

QUERY
Embed → Search → Retrieve → Generate

If you understand these two flows, you understand the foundation of RAG.

Freshness

35. But what happens when documents change?

Suppose on Monday your runbook says "Restart the nginx deployment." We generate an embedding and store it.

On Tuesday, someone updates the runbook to also clear the Redis cache. But imagine our vector database still contains the old version.

The RAG system retrieves outdated instructions. The LLM is not necessarily the problem.

Our index is stale.

Re-index

36. Why embeddings need to be updated

Whenever source documents change, the searchable representation should also be updated appropriately.

GitHub commit
Confluence page updated
New PDF uploaded
S3 object changed
Runbook modified
New incident report created
Document Changed
       ↓
Detect Change
       ↓
Load Document
       ↓
Chunk
       ↓
Generate Embeddings
       ↓
Update Vector Database
DevOps analogy

37. Think of RAG ingestion like CI/CD

A CI/CD pipeline might look like: Code Change → CI Pipeline → Build → Test → Deploy.

A RAG ingestion pipeline can look like:

Document Change
      ↓
Detect Change
      ↓
Load → Chunk → Embed
      ↓
Update Vector Database

Both are automation pipelines. The goal is similar: keep the production system synchronized with the latest source.

Event-driven

38. Event-driven index updates

Git Push
   ↓
GitHub Webhook
   ↓
Ingestion Pipeline
   ↓
Document Loader → Chunk → Embedding
   ↓
Update Vector Database

This can keep the index relatively fresh.

Scheduled

39. Scheduled updates

Not every knowledge base needs immediate updates. You might run hourly, nightly, daily, or weekly:

CronJob
   ↓
Check changed documents
   ↓
Re-index changed content

This is simpler but introduces a delay between document updated and document searchable.

Production

40. Production RAG is more than embeddings

A demo RAG system might be: PDF → Chunks → Embeddings → Vector DB → LLM.

A production RAG system has many more concerns for a DevOps or SRE engineer:

What if the vector database goes down?
What if the embedding API fails?
What if the LLM times out?
What if retrieval returns nothing?
What if the index becomes stale?
What if documents are duplicated?
What if parsing fails?
How do we monitor everything?

This is where RAG becomes an infrastructure problem as much as an AI problem.

Failure mode

41. Failure: vector database is down

Without the vector database, the application cannot retrieve relevant context.

Vector DB unavailable
        ↓
Retry briefly
        ↓
Use replica / secondary endpoint
        ↓
Still unavailable?
        ↓
Return controlled error

Do not pretend retrieval succeeded when it did not.

A safer response might be: "The knowledge search service is temporarily unavailable. Please try again later."

Failure mode

42. Failure: embedding service is down

If the embedding service fails, we cannot create the query vector. Therefore similarity search cannot begin.

Embedding fails → Retry with backoff → Compatible fallback → Controlled failure
Compatibility

43. Be careful with fallback embedding models

Suppose the original index was created using an embedding model that generates 1536-dimensional vectors, but your fallback model generates 768-dimensional vectors.

You cannot simply send that query vector to the existing 1536-dimensional index.

A fallback embedding model must be compatible with the existing index or have its own corresponding index.

Graceful degradation

44. Failure: LLM times out

Retrieval can work perfectly and the LLM can still fail.

Instead of pretending everything failed, the application could return the retrieved sources:

I found relevant runbooks, but the answer-generation service is currently unavailable.

This is an example of graceful degradation.

No results

45. Failure: retrieval finds nothing

This is not necessarily an infrastructure failure. Possible reasons include: knowledge base does not contain the answer, similarity threshold is too strict, unfamiliar terminology, authorization filtering, or incomplete index.

The system should not invent information.

I could not find relevant information in the authorized knowledge base.
Silent failure

46. Failure: stale index

Consider: runbook updated in GitHub → webhook fails → vector database still contains old chunks → RAG retrieves old instructions.

This is dangerous because the system appears healthy. Everything technically works. But the answer is stale.

Useful controls include tracking document version, checksum, source modification time, last indexed time, and synchronization lag — for example index_freshness_seconds.

Duplicates

47. Duplicate documents

Imagine the same runbook exists in GitHub, Confluence, and an old PDF export. Similarity search might return three nearly identical Top-K results.

This wastes context-window space and may cause the LLM to overemphasize repeated information.

Possible controls: document IDs, checksums, content hashes, version metadata, canonical source IDs, and deduplication.

Ingestion failures

48. Failure: document parsing

Parsing may fail because of corrupted PDFs, scanned documents, complex tables, unsupported formats, huge files, or invalid character encoding.

Do not silently ignore these documents.

Document → Parsing fails → Dead-letter queue → Record failure → Retry if appropriate → Alert
Retries

49. Retry strategy

Temporary failures (network errors, HTTP 429, HTTP 503, short timeouts) can often be retried.

Permanent failures (invalid credentials, unsupported file, wrong vector dimensions, malformed request, authorization denied) should not be retried endlessly.

A common strategy is exponential backoff with a maximum, otherwise an outage can create a retry storm.

Protection

50. Circuit breakers

Without protection, thousands of requests can keep hitting an already failing LLM provider.

Failures exceed threshold
          ↓
Circuit opens
          ↓
Requests fail fast or use fallback
          ↓
Provider checked periodically
          ↓
Circuit closes after recovery
Kubernetes

51. Health checks

Liveness answers: is the application process alive? If not, Kubernetes may restart it.

Readiness answers: can this application currently accept traffic?

The RAG API process may be alive while the vector DB is unavailable — alive but not ready.

Do not make every external dependency part of a liveness check, or an external LLM outage can cause endless Pod restarts that do not solve the real problem.

Metrics

52. Observability for RAG

Useful metrics include:

rag_requests_total
rag_request_errors_total
retrieval_latency_seconds
embedding_latency_seconds
llm_latency_seconds
empty_retrieval_total
index_freshness_seconds
chunking_failures_total
vector_db_errors_total
llm_timeouts_total
fallback_model_usage_total
Plan failures

53. Graceful degradation

A production system should decide what happens before failures occur.

Vector DB unavailable → Return controlled service error
LLM unavailable → Return retrieved sources
Primary LLM unavailable → Use secondary LLM
Cache unavailable → Continue without cache
No relevant chunks → Tell user nothing relevant was found
Ingestion delayed → Use existing index with freshness warning
Failure should be predictable and controlled.
Security

54. Security must not disappear during failure

Imagine your normal retrieval flow applies authorization.

A terrible fallback would be: authorization failed → search everything.

A fallback should never weaken security.

If secure retrieval cannot be performed, it is better to fail safely.

Full picture

55. Production RAG architecture

Production RAG architecture with offline indexing into a vector database and an online query path through API gateway, auth, orchestrator, retrieval, prompt builder, LLM, and guardrails.
Ingest once · Retrieve relevant context per query · Generate with the LLM · Validate before returning.

This is much closer to how a DevOps or SRE engineer should think about RAG.

It is not simply "LLM + Vector DB." It is a distributed application with multiple dependencies and failure modes.

Common question

56. RAG vs fine-tuning

One common beginner question is: why not just fine-tune the model with our documents?

These solve different problems.

With RAG, knowledge remains outside the model. At query time we search, retrieve, and give information to the LLM.

This is particularly useful for knowledge that changes frequently: runbooks, incident reports, product documentation, internal procedures, infrastructure documentation.

When the source changes, the retrieval index can be updated. You don't necessarily need to retrain the underlying LLM simply because a runbook changed.

Mental model

57. The most important RAG mental model

                 PREPARE KNOWLEDGE

Documents → Load → Chunk → Embed → Store


                 ANSWER QUESTION

Question → Embed → Search → Retrieve → Add Context → LLM → Answer

That is RAG. Everything else is an improvement, optimization, or production concern around this basic flow.

Before you go

58. If you remember only eight things

1. LLMs do not automatically know your private or newly created company information.

2. RAG lets us retrieve relevant external information before asking the LLM to answer.

3. Documents are loaded and split into smaller chunks.

4. An embedding model converts those chunks into vectors representing their meaning.

5. A vector database stores those vectors and lets us search for semantically similar chunks.

6. The user's question is also converted into an embedding and used for similarity search.

7. Top-K retrieval selects a small number of the most relevant chunks and sends them to the LLM as context.

8. A production RAG system also needs freshness, security, retries, monitoring, failure handling, and graceful degradation.

The shortest possible version is:

Documents → Chunk → Embed → Store

User Question → Embed → Search → Retrieve → Add Context → LLM → Answer

And that is the heart of Retrieval-Augmented Generation.