Embeddings and Vector Databases
What are Embeddings?
An embedding is a numerical representation of text (or other data) as a vector -- a list of floating-point numbers. These numbers capture the semantic meaning of the text, so that texts with similar meanings have vectors that are close to each other in mathematical space.
For example, the sentences "How do I reset my password?" and "I forgot my login credentials" would have very similar vector representations, even though they share almost no words. The embedding model understands that both sentences express the same intent.
This is the fundamental insight that makes semantic search possible: instead of matching keywords, you match meaning.
How Embedding Models Work
Embedding models are neural networks trained to convert text into vectors. The process:
- The text is tokenised (split into subword units)
- Each token is converted to an initial representation
- The model processes all tokens through multiple transformer layers
- The final representation (embedding) is a vector of typically 768 to 3,072 dimensions
Popular embedding models:
- OpenAI text-embedding-3-small: 1,536 dimensions, fast and cheap
- OpenAI text-embedding-3-large: 3,072 dimensions, highest quality
- Cohere Embed: Strong multilingual support
- Sentence-BERT (open-source): Free to run locally
// Run in Node.js -- generating an embedding
import OpenAI from 'openai';
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
async function embedText(text) {
const response = await client.embeddings.create({
model: 'text-embedding-3-small',
input: text,
});
return response.data[0].embedding; // Array of 1536 floats
}
const vector = await embedText('How do I reset my password?');
console.log('Embedding dimensions:', vector.length); // 1536
console.log('First 5 values:', vector.slice(0, 5));
// e.g. [0.0234, -0.1891, 0.0567, 0.2341, -0.0891, ...]
Cosine Similarity: Measuring Semantic Distance
To find documents similar to a query, you measure the similarity between their embeddings. The most common measure is cosine similarity:
- Score of 1.0: identical meaning
- Score of 0.8-0.9: very similar
- Score of 0.5-0.7: somewhat related
- Score below 0.5: likely unrelated
function cosineSimilarity(vecA, vecB) {
const dotProduct = vecA.reduce((sum, a, i) => sum + a * vecB[i], 0);
const magnitudeA = Math.sqrt(vecA.reduce((sum, a) => sum + a * a, 0));
const magnitudeB = Math.sqrt(vecB.reduce((sum, b) => sum + b * b, 0));
return dotProduct / (magnitudeA * magnitudeB);
}
In practice, vector databases calculate this for you -- much faster than doing it in code.
What is a Vector Database?
A vector database is a specialised database designed to store, index, and search vector embeddings at scale. Unlike traditional databases that search by exact match, vector databases search by similarity.
When you store a million document embeddings in a vector database and query with a new embedding, the database returns the top N most similar vectors in milliseconds.
Popular Vector Databases
Pinecone: Fully managed, easy to start, excellent performance. Free tier available. Best for quick starts.
Weaviate: Open-source, self-hostable, built-in multi-modal support. Strong for production.
Chroma: Open-source, lightweight, runs locally. Ideal for development and small projects.
pgvector: PostgreSQL extension that adds vector search to a standard database. Great if you already use PostgreSQL.
Supabase (with pgvector): Managed PostgreSQL with vector search. Familiar to web developers.
Qdrant: Open-source, high performance, rust-based. Excellent for large-scale production.
Building a Simple Embedding Pipeline
// Run in Node.js -- simple RAG pipeline
import OpenAI from 'openai';
import { ChromaClient } from 'chromadb';
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const chroma = new ChromaClient();
async function buildKnowledgeBase(documents) {
const collection = await chroma.createCollection({ name: 'company_docs' });
for (const doc of documents) {
// Generate embedding for each document chunk
const embeddingResponse = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: doc.text,
});
await collection.add({
ids: [doc.id],
embeddings: [embeddingResponse.data[0].embedding],
documents: [doc.text],
metadatas: [{ source: doc.source, title: doc.title }],
});
}
console.log('Knowledge base built with', documents.length, 'documents');
return collection;
}
async function searchKnowledgeBase(collection, query, topK = 5) {
// Embed the query
const queryEmbedding = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: query,
});
// Search for similar documents
const results = await collection.query({
queryEmbeddings: [queryEmbedding.data[0].embedding],
nResults: topK,
});
return results.documents[0].map((doc, i) => ({
text: doc,
metadata: results.metadatas[0][i],
similarity: 1 - results.distances[0][i], // Convert distance to similarity
}));
}
Chunking Strategy
How you split documents into chunks dramatically affects retrieval quality:
Too small (sentence-level): Loses context; retrieved chunks may be too short to be useful.
Too large (full document): Retrieved chunks contain irrelevant content; the model must process more tokens.
Optimal (200-500 words with overlap): Captures meaningful context while staying focused.
function chunkText(text, chunkSize = 400, overlap = 50) {
const words = text.split(' ');
const chunks = [];
for (let i = 0; i < words.length; i += chunkSize - overlap) {
const chunk = words.slice(i, i + chunkSize).join(' ');
if (chunk.split(' ').length >= 50) { // Minimum chunk size
chunks.push(chunk);
}
}
return chunks;
}
Metadata Filtering
Vector databases support filtering by metadata alongside semantic search:
// Search only within a specific document category
const results = await collection.query({
queryEmbeddings: [queryEmbedding],
nResults: 5,
where: { category: 'refund_policy' }, // Metadata filter
});
// Combine semantic search with date filtering
const recentResults = await collection.query({
queryEmbeddings: [queryEmbedding],
nResults: 5,
where: { updated_after: '2025-01-01' },
});
Metadata filtering is essential in large knowledge bases where you only want to search within specific document sets.
Key Takeaways
- An embedding is a numerical vector representation of text that captures semantic meaning -- similar texts have similar vectors.
- Cosine similarity measures how semantically close two embeddings are, enabling meaning-based document retrieval.
- Vector databases store and index embeddings for fast similarity search across millions of documents.
- Chunking strategy (200-500 words with overlap) significantly affects retrieval quality -- too small loses context, too large adds noise.
- Metadata filtering allows you to combine semantic search with structured filters (category, date, author) for precise retrieval.
Try it yourself
Key Takeaways
- Embeddings are numerical vector representations of text that capture semantic meaning -- similar texts produce similar vectors.
- Cosine similarity measures semantic closeness between embeddings -- the primary metric for retrieval in RAG systems.
- Vector databases (Pinecone, Chroma, Weaviate, pgvector) store and index embeddings for fast similarity search at scale.
- Optimal chunk size (200-500 words with overlap) balances context richness with retrieval precision.
- Metadata filtering combines vector search with structured constraints, enabling precise retrieval within large knowledge bases.
Quick Quiz
1.What is an embedding in the context of AI and NLP?
2.What does cosine similarity measure?
3.What chunking size typically works best for RAG document retrieval?
4.What is the advantage of using pgvector over a dedicated vector database like Pinecone?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx