RAG Best Practices and Optimisation
Moving from Working to Excellent
A basic RAG system is not hard to build. Getting it to work reliably, accurately, and efficiently at production scale is where the real engineering happens. This lesson covers the optimisation techniques that separate a production RAG system from a prototype.
Optimisation 1: Better Chunking
The default approach of splitting by word count is a starting point, not an end point.
Semantic chunking
Split documents at natural boundaries (paragraphs, sections, topic transitions) rather than arbitrary word counts:
function semanticChunk(text) {
// Split on double newlines (paragraph boundaries) first
const paragraphs = text.split(/\n\n+/);
const chunks = [];
let currentChunk = '';
for (const para of paragraphs) {
if ((currentChunk + para).split(' ').length > 500) {
if (currentChunk) chunks.push(currentChunk.trim());
currentChunk = para;
} else {
currentChunk += (currentChunk ? '\n\n' : '') + para;
}
}
if (currentChunk) chunks.push(currentChunk.trim());
return chunks.filter(c => c.split(' ').length >= 30);
}
Hierarchical chunking
Store both small chunks (for precise retrieval) and larger parent chunks (for full context). Retrieve small chunks, but pass the parent chunk to the LLM:
Small chunk retrieved: "Refunds are processed within 5-7 business days."
Parent chunk passed to LLM: Full refund policy section (400 words)
This gives you the precision of small-chunk retrieval with the context richness of larger chunks.
Optimisation 2: Query Transformation
Users ask questions in many different ways. Transform the query before retrieval to improve coverage:
Query expansion
Generate multiple variations of the query and retrieve for all of them:
async function expandQuery(query) {
const response = await openai.chat.completions.create({
model: 'gpt-4o-mini',
temperature: 0,
messages: [{
role: 'user',
content: 'Generate 3 alternative phrasings of this question for a search engine. Return as JSON array: ["q1", "q2", "q3"]. Question: ' + query
}],
response_format: { type: 'json_object' },
});
const alternatives = JSON.parse(response.choices[0].message.content).alternatives;
return [query, ...alternatives];
}
// Retrieve for all query variations, deduplicate results
const queries = await expandQuery("How do I get a refund?");
// Might generate: ["How do I get a refund?", "What is your return policy?", "Can I get my money back?", "Refund process steps"]
Hypothetical Document Embedding (HyDE)
Generate a hypothetical answer to the query, then retrieve documents similar to that hypothetical answer rather than the query itself. This often retrieves more relevant documents because the hypothetical answer is in the same "vocabulary space" as the actual documents.
Optimisation 3: Hybrid Search
Pure semantic search misses documents with exact keyword matches. Pure keyword search misses semantic variations. Hybrid search combines both:
async function hybridSearch(collection, query, topK = 10) {
// Semantic search
const queryEmbedding = await embedText(query);
const semanticResults = await collection.query({
queryEmbeddings: [queryEmbedding],
nResults: topK,
});
// Keyword search (using BM25 or simple text matching)
const keywordResults = await collection.query({
queryTexts: [query], // ChromaDB also supports text-based search
nResults: topK,
});
// Merge and re-rank results using Reciprocal Rank Fusion (RRF)
return rerankRRF(semanticResults, keywordResults);
}
function rerankRRF(semanticResults, keywordResults, k = 60) {
const scores = {};
semanticResults.ids[0].forEach((id, rank) => {
scores[id] = (scores[id] || 0) + 1 / (k + rank + 1);
});
keywordResults.ids[0].forEach((id, rank) => {
scores[id] = (scores[id] || 0) + 1 / (k + rank + 1);
});
return Object.entries(scores)
.sort(([, a], [, b]) => b - a)
.map(([id, score]) => ({ id, score }));
}
Optimisation 4: Re-ranking
After retrieving 10-20 candidate chunks, use a dedicated re-ranking model to select the top 3-5:
async function rerankWithLLM(query, chunks, topN = 3) {
const response = await openai.chat.completions.create({
model: 'gpt-4o-mini',
temperature: 0,
messages: [{
role: 'user',
content: `Rank these document chunks by relevance to the query. Return JSON: { "ranked_indices": [most_relevant, ..., least_relevant] }
Query: ${query}
Chunks:
${chunks.map((c, i) => '[' + i + '] ' + c.text.substring(0, 200)).join('\n\n')}`
}],
response_format: { type: 'json_object' },
});
const ranked = JSON.parse(response.choices[0].message.content).ranked_indices;
return ranked.slice(0, topN).map(i => chunks[i]);
}
For production use, dedicated cross-encoder models (like Cohere Rerank or BGE-Reranker) are faster and cheaper than using an LLM for re-ranking.
Optimisation 5: Caching
Two caching strategies significantly reduce cost and latency:
Query caching
Cache the results of recent queries. If the same (or very similar) question is asked again, return the cached answer:
const queryCache = new Map();
const CACHE_TTL = 3600 * 1000; // 1 hour in milliseconds
async function cachedAnswer(query, collection) {
const cacheKey = query.toLowerCase().trim();
const cached = queryCache.get(cacheKey);
if (cached && Date.now() - cached.timestamp < CACHE_TTL) {
return { ...cached.result, fromCache: true };
}
const result = await answerQuestion(query, collection);
queryCache.set(cacheKey, { result, timestamp: Date.now() });
return result;
}
Embedding caching
The same document chunk will always produce the same embedding. Cache embeddings to avoid regenerating them:
const embeddingCache = new Map();
async function cachedEmbed(text) {
if (embeddingCache.has(text)) return embeddingCache.get(text);
const embedding = await embedText(text);
embeddingCache.set(text, embedding);
return embedding;
}
Optimisation 6: Source Citation and Trust
In production, users need to trust the system's answers. Always cite sources:
// Enhanced response format with citations
async function generateAnswerWithCitations(query, chunks) {
const response = await openai.chat.completions.create({
model: 'gpt-4o-mini',
temperature: 0,
messages: [
{ role: 'system', content: 'Answer using only the provided sources. End every claim with [Source N] where N is the source number.' },
{ role: 'user', content: 'Sources:\n' + chunks.map((c, i) => '[Source ' + (i+1) + ': ' + c.metadata.source + ']\n' + c.text).join('\n\n') + '\n\nQuestion: ' + query }
],
});
return {
answer: response.choices[0].message.content,
sources: chunks.map(c => ({ title: c.metadata.source, url: c.metadata.url })),
};
}
Monitoring RAG in Production
Track these metrics continuously:
- Retrieval precision: What percentage of retrieved chunks are actually relevant?
- Answer accuracy: What percentage of answers are factually correct (sample and evaluate)?
- No-answer rate: How often does the system say "I don't know"? Too high suggests poor coverage; too low suggests hallucination.
- Embedding freshness: How often are new documents ingested? Stale knowledge bases give outdated answers.
- Latency: p50 and p95 end-to-end response time (target: under 2 seconds for most use cases)
Key Takeaways
- Semantic and hierarchical chunking outperform fixed-size chunking by respecting document structure and providing rich context.
- Query expansion (generating alternative phrasings) significantly improves retrieval coverage for diverse user language.
- Hybrid search (semantic + keyword) consistently outperforms either approach alone for real-world queries.
- Re-ranking retrieved chunks with a cross-encoder model before passing to the LLM improves answer quality noticeably.
- Caching query results and embeddings reduces latency and cost significantly for high-volume systems.
Try it yourself
Key Takeaways
- Semantic chunking (splitting at paragraph boundaries) and hierarchical chunking (small retrieve, large context) outperform fixed word-count chunking.
- Query expansion generates alternative phrasings before retrieval, improving coverage for the full range of user language.
- Hybrid search combines semantic vector search with keyword search, consistently outperforming either approach alone.
- Re-ranking retrieved candidates with a cross-encoder model before passing to the LLM noticeably improves answer quality.
- Monitor retrieval precision, answer accuracy, no-answer rate, and latency continuously in production to detect and fix regressions.
Quick Quiz
1.What is hierarchical chunking and why is it useful?
2.What is Hypothetical Document Embedding (HyDE)?
3.Why is caching particularly valuable in a RAG system?
4.What does 'no-answer rate' measure in a RAG monitoring system?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx