LLM Engineering8 min read
Retriever k-value tuning for RAG: the right top-k
How to pick the right k value for your RAG retriever. The 3-step tuning process, the failure modes of k=3 and k=20, and the sweet spot in between.
Build Your First Software Factory Execution Harness live Thursday, Oct 1, 12:00 PM ET
Oct 1, Reserve a seatExplore our latest articles and insights about Vector Databases.
10 posts in total
LLM Engineering8 min read
How to pick the right k value for your RAG retriever. The 3-step tuning process, the failure modes of k=3 and k=20, and the sweet spot in between.
LLM Engineering8 min read
How to combine multiple vector stores in one RAG pipeline. The merge pattern, the deduplication rule, and when multi-source beats a single index.
LLM Engineering8 min read
How to use FAISS for production RAG. Index types, persistence, memory trade-offs, and the 4 settings that decide if FAISS beats a managed vector DB.
LLM Engineering11 min read
How to pick an embedding model for production RAG. The 5 criteria that matter, the benchmarks that lie, and the migration cost nobody warns you about.
LLM Engineering11 min read
How hybrid retrieval combines vector search and graph traversal in RAG. The when, the why, and the 60-line fusion that beats either alone.
AI Engineering in Practice10 min read
Understand how embeddings and vector databases work under the hood. Learn how computers translate text meaning into numbers and search millions of docum...
AI Engineering13 min read
Compare vector databases for RAG systems. Learn when to use Chroma, Qdrant, Weaviate, pgvector, Pinecone, and Vespa based on performance, scale, and dev...
AI Engineering8 min read
Learn how embeddings and vector databases power RAG systems. Understand semantic search, cosine similarity, metadata filtering, and choose between open-...
AI Engineering7 min read
Master document chunking for RAG systems. Learn fixed-size, recursive, semantic, and content-aware splitting techniques to improve retrieval quality and...
AI Engineering9 min read
How retrieval works in production. Learn how RAG solves LLM limitations by connecting models to external documents. Master chunking, embeddings, vector dat
10 posts in total
LLM Engineering8 min read
How to pick the right k value for your RAG retriever. The 3-step tuning process, the failure modes of k=3 and k=20, and the sweet spot in between.
LLM Engineering8 min read
How to combine multiple vector stores in one RAG pipeline. The merge pattern, the deduplication rule, and when multi-source beats a single index.
LLM Engineering8 min read
How to use FAISS for production RAG. Index types, persistence, memory trade-offs, and the 4 settings that decide if FAISS beats a managed vector DB.
LLM Engineering11 min read
How to pick an embedding model for production RAG. The 5 criteria that matter, the benchmarks that lie, and the migration cost nobody warns you about.
LLM Engineering11 min read
How hybrid retrieval combines vector search and graph traversal in RAG. The when, the why, and the 60-line fusion that beats either alone.
AI Engineering in Practice10 min read
Understand how embeddings and vector databases work under the hood. Learn how computers translate text meaning into numbers and search millions of docum...
AI Engineering13 min read
Compare vector databases for RAG systems. Learn when to use Chroma, Qdrant, Weaviate, pgvector, Pinecone, and Vespa based on performance, scale, and dev...
AI Engineering8 min read
Learn how embeddings and vector databases power RAG systems. Understand semantic search, cosine similarity, metadata filtering, and choose between open-...
AI Engineering7 min read
Master document chunking for RAG systems. Learn fixed-size, recursive, semantic, and content-aware splitting techniques to improve retrieval quality and...
AI Engineering9 min read
How retrieval works in production. Learn how RAG solves LLM limitations by connecting models to external documents. Master chunking, embeddings, vector dat