Data Intelligence · RAG · PostgreSQL · Semantic Search

pgvector recall is a setting you chose: baseline, candidate list, measurement

Published · Updated

pgvector's default is exact search with perfect recall, so lost recall is a trade someone made. Field notes on the baseline, the candidate list and the measurement to run before blaming the embeddings.

Teams usually describe pgvector recall as something that degrades: answers were good at first, then the corpus grew and retrieval went soft. The documentation describes it differently. By default, pgvector performs exact nearest neighbour search, which provides perfect recall. An index is what you add to use approximate nearest neighbour search, and it trades some recall for speed. Nothing drifts on its own. A trade was made, usually once, usually without a baseline.

Start from the ceiling you already have

That framing changes the debugging order. Before asking what broke, re-run the same queries with no index and compare the result sets. Exact search gives you the ceiling, and that ceiling stays cheap to compute even when it is far too slow to serve in production. Any gap between the two sets is the price of the index as currently configured, not evidence that your embeddings stopped working.

pgvector gives you distinct choices inside that trade. An HNSW index creates a multilayer graph and has better query performance than IVFFlat in terms of the speed-recall trade-off, but it has slower build times and uses more memory. On the search side you specify the size of the dynamic candidate list, which is 40 by default. If nobody has touched that value, you are serving whatever recall the default produces on your data.

Pick a metric before you tune anything

Anthropic publishes a metric worth copying: 1 minus recall@20, which measures the percentage of relevant documents that fail to be retrieved within the top 20 chunks. It needs a fixed question set with known-correct chunks, and building that set is the step most teams skip. Without it, every tuning change is argued from anecdote, and a change that rescues one demo query while quietly hurting the rest of the corpus will look like progress.

Some failures no index can fix

One class of retrieval failure survives all parameter tuning. Embedding models excel at capturing semantic relationships, but they can miss crucial exact matches, and real catalogues are full of those: part references, error codes, model names, strings that differ by a single character. No candidate list size recovers a token the vector never encoded distinctly. What helps instead is a second retrieval path running over the same rows.

Postgres already ships that second path

Full-text searching provides the capability to identify natural-language documents that satisfy a query, and optionally to sort them by relevance to the query. The tsvector type stores preprocessed documents and tsquery represents processed queries, while full-text indexing allows documents to be preprocessed and an index saved for later rapid searching. The documentation is blunt about the shortcut many teams reach for first: the ~, ~*, LIKE and ILIKE operators lack many essential properties required by modern information systems.

Content and ranking outrank parameters

Index settings are rarely the highest-leverage knob either. Contextual retrieval, which attaches surrounding context to chunks before they are embedded, can reduce the number of failed retrievals by 49%, and by 67% when combined with reranking. Those are changes to what you store and how you rank it, not to how hard you search. Work in that order: chunk context first, reranking second, candidate list last, and measure between each step.

Check whether you need retrieval at all

There is also an exit worth testing early. The guidance is explicit: if your knowledge base is smaller than 200,000 tokens, about 500 pages of material, you can include the entire knowledge base in the prompt you give the model, with no need for RAG or similar methods. Plenty of internal corpora are smaller than that and have been wrapped in an index nobody needed, along with the tuning work that follows.

A corpus that grows needs a repeated measurement

Any of this ages. On Tatano Energy we run a multilingual platform across 4 country domains indexed separately, serving 7 languages, publishing 8 SEO articles every day with 0 manual intervention required. A corpus that grows on a schedule is not the corpus you measured at launch. Whatever recall your settings deliver today was measured against yesterday's rows, which is an argument for re-running the evaluation set rather than re-tuning on instinct.

Treat recall as a configured value with a known ceiling, not a mystery. Keep an exact-search baseline you can re-run, keep a fixed question set and a metric like 1 minus recall@20, and change one thing at a time: chunk context, then reranking, then the candidate list. If the corpus grows on a schedule, put the evaluation on the same schedule. Most arguments about vector search are really arguments about missing measurement.

Sources

pgvector (GitHub) — Open-source vector similarity search for Postgres — https://github.com/pgvector/pgvector

Anthropic — Introducing Contextual Retrieval — https://www.anthropic.com/engineering/contextual-retrieval

PostgreSQL 18 Documentation — Full Text Search: 12.1. Introduction — https://www.postgresql.org/docs/current/textsearch-intro.html

Neurolinks case study — Four markets, one codebase — https://neurolinks.be/work/tatano-energy

Working on a project where these methods apply?