A vector database stores embeddings, which are lists of numbers that represent text, images or other data, and returns the stored vectors closest to a query vector. That sounds like a product category, but the hard part is one question: do you compare the query with every vector, or only with some of them? This guide explains vector databases through that choice, because it determines cost, latency and the failure modes you will meet. It builds on how RAG retrieval works, where embeddings are the first stage, and on tokens, context windows and AI cost.
What a vector database actually does
Two jobs: it stores vectors alongside identifiers and metadata, and it answers nearest-neighbour queries under a distance measure. pgvector, an open-source Postgres extension, lists the common measures as operators: <-> for L2 (Euclidean) distance, <#> for negative inner product, <=> for cosine distance, and <+> for L1 distance, with Hamming and Jaccard for bit vectors.[2] The measure should match the one your embedding model was trained for; check the model's documentation.
Everything else (APIs, sharding, replication, hosted dashboards) varies by product. The part that is common to all of them is the index, and that is where the trade-offs sit.
Exact search: the baseline you should keep
The simplest index is no clever index at all. A "flat" index stores the vectors in an array and compares the query with all of them. The Faiss documentation marks its flat indexes as exhaustive and gives their memory use as 4*d bytes per vector, which is four bytes for each of the d float32 dimensions.[3] pgvector behaves the same way until you add an index: "By default, pgvector performs exact nearest neighbor search, which provides perfect recall."[2]
Exact search has two attractions. Results are the true nearest neighbours, so there is nothing to tune and nothing to explain. And it is the ground truth you need to measure any approximate index against.
The cost is that query time grows with the number of vectors. As arithmetic from the Faiss formula (our calculation, not a benchmark): at 1,024 dimensions, one flat float32 vector takes 4,096 bytes, so one million vectors need roughly 4.1 GB and one hundred million roughly 410 GB, before any metadata or index overhead. Whether brute force is fast enough at your size is something to measure on your hardware, not assume.
Approximate search: skipping most of the data
Approximate nearest neighbour (ANN) indexes avoid comparing the query with everything. pgvector states the bargain plainly: an approximate index "trades some recall for speed," and "unlike typical indexes, you will see different results for queries after adding an approximate index."[2] That second sentence matters operationally: adding an index can change your answers, so it needs testing like a code change.
Recall here means the share of the true nearest neighbours that the approximate search returns. Two designs dominate.
HNSW: a layered graph
Hierarchical Navigable Small World (HNSW) was introduced by Malkov and Yashunin. It builds a multi-layer structure of proximity graphs over nested subsets of the data. Each element's top layer is chosen randomly with an exponentially decaying probability, and a search starts in the sparse top layer and moves down. The authors say this combination "allows a logarithmic complexity scaling" and describe the method as fully graph-based, with no additional search structures.[1]
In practice you get the following from pgvector's documentation:[2]
- HNSW has better speed-recall performance than IVFFlat, but builds more slowly and uses more memory.
- It can be created on an empty table, because there is no training step.
- Its main settings and documented defaults are
m(maximum connections per layer, default 16),ef_construction(default 64, higher gives better recall at the cost of build time) and the query-timehnsw.ef_search(default 40, higher gives better recall at the cost of speed).
The Faiss wiki gives the memory formula for its HNSW index as 4*d + x * M * 2 * 4 bytes per vector, and notes that a larger M is more accurate but uses more memory.[3] The graph is extra memory on top of the vectors themselves.
IVF: clusters you search selectively
An inverted-file (IVF) index first clusters the vectors, then searches only the clusters closest to the query. pgvector describes IVFFlat as dividing vectors into lists and searching "a subset of those lists that are closest to the query vector." It builds faster and uses less memory than HNSW, but has lower query performance at a given recall.[2] Unlike HNSW it needs a training step, because the cluster centroids are learned from the data (in Faiss, the IVF train method adds the centroids to a coarse index). The Faiss documentation names the failure case: "when the cell of the nearest neighbor of a given query is not selected."[3]
The tuning setting is how many clusters to visit. In pgvector, ivfflat.probes defaults to 1, can be set to the number of lists to get exact search, and sqrt(lists) is suggested as a starting point. For the number of lists, the README suggests rows / 1000 up to one million rows and sqrt(rows) above that.[2]
Compression for very large collections
Faiss also offers quantisation: scalar quantisation (SQ8 needs about d bytes per vector instead of 4*d) and product quantisation, where memory is ceil(M * nbits / 8) bytes. Its IVF-PQ index is described on the wiki as "probably the most useful index for large-scale search".[3] Compression buys memory at the price of accuracy, so the same rule applies: compare against exact results. For scale, a 2017 paper by Johnson, Douze and Jégou reports building a k-NN graph over 1 billion vectors in under 12 hours on four GPUs, which shows what specialised GPU implementations can reach and why such systems are a different class from a Postgres extension.[4]
How the options compare
| Option | Recall | Memory | Build and tuning | Source of the facts |
|---|---|---|---|---|
| Flat / exact | Perfect | 4*d bytes per vector | Nothing to tune | Faiss wiki, pgvector README |
| HNSW | Approximate | Vectors plus graph links | Slower build; m, ef_construction, ef_search | pgvector README, Faiss wiki, HNSW paper |
| IVFFlat | Approximate | Vectors plus 8 bytes each (Faiss) | Faster build; needs training data; lists, probes | pgvector README, Faiss wiki |
| IVF-PQ | Approximate, compressed | Far below flat | Training; more parameters | Faiss wiki |
The rows state documented properties, not measured speeds. Speed and recall depend on data, dimensions and hardware.
The pitfall: filters and approximate indexes
Real queries rarely say only "nearest to this vector". They say "nearest to this vector among documents this user may read, from this year". pgvector's documentation warns that with approximate indexes "filtering is applied after the index is scanned." A selective filter can therefore leave fewer results than LIMIT requests. The README's example: with a filter matching 10% of rows and the default hnsw.ef_search of 40, only about 4 rows will match on average.[2]
The documented mitigations are iterative index scans (available from pgvector 0.8.0), which keep scanning until enough results are found, a choice between strict and relaxed ordering, partial indexes for filters with few distinct values, and partitioning for many. Exact search remains a good option when the filter matches a small share of rows.[2] Other products handle filtering differently; this is the question to put to any vendor.
When plain Postgres is enough, and when it is not
This is our interpretation of the documented trade-offs, not a measured result:
- Start with exact search if the corpus is small, and keep it as a test oracle. If it meets your latency target, you have nothing to maintain.
- Add HNSW when exact search is too slow and memory allows. It is the documented better speed-recall option, and needs no training step.
- Consider IVF or compressed indexes when memory or build time is the binding limit, accepting lower recall at the same speed.
- Look beyond one machine when the corpus no longer fits in memory, or you need GPU throughput or distributed serving. That is where dedicated systems earn their place.
Whichever route you take, build a small evaluation set: queries with known relevant passages, an exact-search baseline, and recall measured as you change parameters. Our retrieval guide describes why retrieval should be measured separately from answer quality.
What this article does not cover
We did not benchmark any product, so no speed or recall numbers appear here beyond those quoted from the sources. Defaults and limits come from pgvector and Faiss documentation as read on 2026-10-09 and change between versions; check the current documentation before relying on them. Hosted vector databases differ from these libraries in features and pricing and are outside this article.




