This section builds the retrieval substrate underneath agent memory: embeddings, indexes, hybrid search, chunking, and the precision/recall trade-offs that decide what gets recalled.
What you will learn
- How embeddings, ANN indexes, and distance metrics work in practice
- When keyword and hybrid search beat pure vector retrieval
- Chunking, reranking, filtering, and freshness for memory workloads
- How embedding drift and query formulation affect recall quality
Chapters in this section
There are 24 chapters in this part. Open any chapter to read it on its own, or work through them in order.
- 1 What Is a Vector Embedding?
- 2 Embedding Models and Semantic Similarity
- 3 Vector Indexes: How Approximate Nearest Neighbor Search Works
- 4 HNSW: The Graph-Based Index Behind Modern Vector Search
- 5 Distance Metrics: Cosine, Dot Product, and Euclidean
- 6 Why Vector Search Alone Is Not Enough for Memory
- 7 Keyword Search and the Inverted Index
- 8 BM25: Scoring Relevance Without Embeddings
- 9 Hybrid Search: Combining Vector and Keyword Retrieval
- 10 The Alpha Parameter: Tuning the Balance Between Semantic and Lexical Search
- 11 Reranking: A Second Pass for Precision
- 12 Late Interaction Retrieval Explained
- 13 Multi-Vector Representations
- 14 Vector Quantization and Compression Trade-offs
- 15 Filtered Vector Search: Combining Metadata and Similarity
- 16 Chunking Strategies and Why They Matter for Memory
- 17 Hierarchical Chunking for Long Documents
- 18 Late Chunking: Preserving Context During Segmentation
- 19 Recall vs Precision in Memory Retrieval
- 20 Approximate vs Exact Search Trade-offs at Scale
- 21 Index Rebuilding and Vector Freshness
- 22 Embedding Drift: When Meaning Shifts Over Time
- 23 Choosing an Embedding Model for a Memory System
- 24 Query Formulation for Effective Memory Recall