What embeddings are, how vector search finds related content by meaning, and how to decide whether you need a dedicated vector database.
In one sentence
An embedding turns a piece of text into a list of numbers so that texts with similar meaning end up close together. Vector search uses that closeness to find relevant content even when the exact words differ.
What an embedding is
An embedding model reads a sentence, paragraph or image and outputs a fixed-length list of numbers, called a vector. The model is trained so that related inputs produce vectors that point in similar directions. "How do I reset my password?" and "I forgot my login" will be close together, even though they share almost no words.
How similarity is measured
To compare two vectors, systems usually calculate cosine similarity or a dot product. A higher score means the two texts are more alike in meaning. A search compares the query vector with stored document vectors and returns the closest ones.
Why indexes are needed
Comparing a query with every stored vector becomes slow at large scale. Vector indexes use approximate nearest-neighbour algorithms, such as HNSW, to find very close matches quickly without checking everything. The trade-off is a small chance of missing the true best match in exchange for much faster search.
Where vectors are stored
Options include dedicated vector databases, vector extensions for existing databases such as PostgreSQL with pgvector, and search engines that support both keyword and vector queries. For many teams, adding vector search to a database they already run is simpler than introducing a new system. Dedicated services become attractive at larger scale or when you need advanced filtering and operations features.
Limits of vector search
Embeddings capture meaning but can miss exact terms such as product codes, names or error numbers. Hybrid search, which combines keyword ranking such as BM25 with vector similarity, usually handles both. Results also depend heavily on the embedding model; one trained for your language and domain will generally perform better.
Choosing an embedding model
Consider language coverage, domain fit, vector size (which affects storage and speed), cost and whether the model can run on your own infrastructure. Public benchmarks such as MTEB are a useful starting point, but test on your own queries. Changing the model later means re-embedding all stored content, so choose deliberately.
Key takeaways
- Embeddings place similar meanings close together as vectors.
- Vector search finds related content without exact keyword matches.
- Hybrid search covers exact terms that embeddings can miss.
- You may not need a new database; many existing ones support vectors.
If you are designing search or retrieval over your own content, our RAG team can help you choose the embedding model and storage that fit.
RAG Development Services- MTEB: Massive Text Embedding BenchmarkMuennighoff et al. · 2022 · Research paperarxiv.org
- Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs (HNSW)Malkov & Yashunin · 2016 · Research paperarxiv.org
- pgvector: open-source vector similarity search for Postgrespgvector project · Documentationgithub.com