Exact matches are a different index: tsvector beside pgvector
Published · Updated
Embedding search misses exact matches by design. The repair is a second retrieval path in the same database, a corpus-size check before building anything, and a metric that counts misses.
A retrieval system that answers conceptual questions well and then fails on a part number, an error code or an invoice reference is not badly tuned. It is doing what it was built to do. Anthropic states the limit plainly: embedding models excel at capturing semantic relationships, but they can miss crucial exact matches. The repair is a second retrieval path, not a larger embedding model.
The failure mode is narrow
Dense retrieval ranks documents by proximity in vector space. A query dominated by an opaque token carries little semantic signal, so the nearest neighbours are documents that merely look topically related. Nothing downstream recovers from that. A reranker reorders a candidate set; it cannot promote a document that never entered the set. The miss happens at the first stage, which is where the fix belongs.
Anthropic reports that contextual retrieval can reduce the number of failed retrievals by 49%, and by 67% when combined with reranking. Read those as reductions, not elimination. A system whose failure rate has been cut still fails on some queries, so the real design question is what happens on a miss: a visible empty result, or a confident answer built on the wrong chunk.
Postgres ships the lexical half
The other half is already installed. The PostgreSQL documentation describes full text search as the capability to identify natural-language documents that satisfy a query, and optionally to sort them by relevance to the query. A tsvector type stores preprocessed documents and a tsquery type represents processed queries. Full text indexing allows documents to be preprocessed and an index saved for later rapid searching.
One warning in those docs is worth taking literally. Postgres has the ~, ~*, LIKE and ILIKE operators for textual data types, but they lack many essential properties required by modern information systems. A pattern match bolted onto a vector query is not a lexical retrieval path. If the second path matters enough to build, give it a real index and a relevance ordering rather than a substring filter.
First check whether you need retrieval at all
Before building two indexes and a fusion step, measure the corpus. Anthropic's guidance is explicit: if your knowledge base is smaller than 200,000 tokens, about 500 pages of material, you can just include the entire knowledge base in the prompt that you give the model, with no need for RAG or similar methods. Many internal knowledge bases sit under that line. An architecture you do not need is still one you have to operate.
Measure misses, not hits
Anthropic uses 1 minus recall@20 as its evaluation metric, which measures the percentage of relevant documents that fail to be retrieved within the top 20 chunks. That framing is useful because it counts misses. Average relevance scores hide the queries that return nothing usable, and those are exactly the queries that produce wrong answers. Fix a query set with known correct documents, then track failures against it.
The vector side remains a configured trade-off
Adding a lexical path does not settle the vector path. By default, pgvector performs exact nearest neighbour search, which provides perfect recall; adding an index switches to approximate search, which trades some recall for speed. An HNSW index creates a multilayer graph with better query performance than IVFFlat on the speed-recall trade-off, at the cost of slower build times and more memory. The dynamic candidate list for search is 40 by default.
Provenance does not come free
The original RAG paper pairs a pre-trained seq2seq model as parametric memory with a dense vector index of Wikipedia as non-parametric memory, accessed with a pre-trained neural retriever. It set the state of the art on three open-domain QA tasks, and the authors found that RAG models generate more specific, diverse and factual language than a parametric-only baseline. They also note that providing provenance for decisions and updating world knowledge remain open research problems.
Treat retrieval as two paths and one measurement. Check first whether the corpus is small enough to skip retrieval entirely. If it is not, give exact matches a tsvector index rather than a LIKE clause, keep the vector path's recall trade-off an explicit choice, and report 1 minus recall@20 separately for identifier queries and conceptual ones. Then decide what the system shows a user when both paths miss.
Sources
pgvector (GitHub) — Open-source vector similarity search for Postgres — https://github.com/pgvector/pgvector
arXiv — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., arXiv:2005.11401) — https://arxiv.org/abs/2005.11401
Anthropic — Introducing Contextual Retrieval — https://www.anthropic.com/engineering/contextual-retrieval
PostgreSQL 18 Documentation — Full Text Search: 12.1. Introduction — https://www.postgresql.org/docs/current/textsearch-intro.html
Working on a project where these methods apply?