Module 03 / 08

Retrieval & RAG

Learn semantic retrieval, chunking, and grounded generation as one system.

The point

Retrieval and generation fail together in real products. Treating them as one layer makes it easier to debug relevance, citations, and answer quality.

Start with

Understanding check

You should be able to…

  • explain semantic search and cosine similarity
  • choose a chunking strategy for a document type
  • inspect retrieved context before blaming the model
  • separate retrieval quality from answer quality
  • create a small eval set for grounded answers

Practice with an agent

Add cited answers to your product

beginner · starter repository, model API, prepared documents

Make

a small retrieval slice that answers from a prepared document set and exposes the retrieved chunks

You know it works when

answer five questions with citations and identify whether each failure came from retrieval or generation

Go deeper by building

Build a cited knowledge assistant

intermediate · next.js or python, embeddings API, sqlite or a vector store, model API

Make

a local app that searches a document folder and answers with citations

You know it works when

include retrieval logs, citations, and ten grounded-answer evals with pass/fail notes