Someone searches a health library for “irregular heartbeat”. The most useful page may be titled “Atrial fibrillation”, with “AF” in the body. A keyword search can miss the different wording. A search based on meaning can find the page, but it might also bring back another heart condition. If the page is old or the reader cannot access it, neither search method has solved the problem.
A vector database helps with that wording gap. It finds material that is similar in meaning to a question. The application still has to decide which material to trust, which person may see it, and whether it is enough to support an answer.
An embedding model turns text into a list of numbers called a vector. You can think of the numbers as coordinates on a map. Passages with related meaning often land near one another, even when they use different words. The model learns that map from its training data; it does not discover medical truth or verify the source.
The database stores the vector alongside an ID that leads back to the original passage. A useful record also carries its source URL, title, date, access rules and the model version that made the vector. Without those links, the search result is just a promising set of numbers.
Here is one of the small, synthetic records in LumenTrail, the project that accompanies this post:
passage = {"id": "demo-af","question": "What is atrial fibrillation?","answer": "Atrial fibrillation, often shortened to AF, is an irregular heart rhythm.","source": "synthetic-demo","entities": ["atrial fibrillation", "AF"],}
The text is deliberately simple and is not medical advice. It lets us inspect what search did without downloading a large dataset first.
When a new question arrives, the application embeds it with the same model and looks for nearby stored vectors. Cosine similarity is a common way to compare their direction. Dot product and Euclidean distance are also used. Follow the embedding model’s guidance rather than switching distance measures casually. A score of 0.8 means the vectors are close under that measure; it does not mean an answer is 80% true.
Before search, I clean each source and split it into passages, often called chunks. A chunk should be small enough to retrieve precisely and large enough to keep a fact with its warning, exception or heading. Splitting a drug instruction from its caution is a poor trade, however tidy the chunk size looks.
There is no chunking method that always finds the right boundary. A fixed character count is an easy baseline, but it may cut a sentence. Splitting around paragraphs or headings keeps more of the source’s structure. A model-based semantic splitter may group related ideas, but it adds a model and can still separate a warning from the fact it qualifies. LumenTrail’s chunking notebook lets you see the first two methods on synthetic text. I would choose among them by checking retrieved evidence on reviewed questions, not by how neat the chunks look.
At question time, I search both the original words and the vectors. Keyword search, often ranked with BM25, is good at exact names, acronyms and identifiers. Dense vector search is good at paraphrases. I merge the candidate lists, optionally use a reranker to take a closer look at the top results, and send a small set of passages to the answer model with source IDs.
source → clean → chunk → record source and access rules → index words and vectorsquestion → check access → keyword + vector candidates → merge → rerank→ cited context → answer, or say that the evidence is missing
The access check must be enforced by application code. Telling a model to respect permissions in a prompt is not an access control system. The same applies to deleted content: a source that has been removed must stop appearing in search, not merely be hidden from the final answer.
LumenTrail makes this route visible. Its notebook builds a five passage index and runs each retrieval method over it:
from lumentrail.retrieval import searchhits = search(database, "What is AF?", embedder, method="hybrid", limit=2)[(hit["id"], hit["source"]) for hit in hits]
The notebook’s offline hash vectors are useful for testing the route, but they do not understand synonyms. That limitation is intentional: you can see the pipeline work before downloading a trained model. The optional BGE embedding shows what real dense retrieval adds.
Consider two questions about the same page: “What is AF?” and “What causes an irregular heartbeat?” The first has an exact acronym. The second uses everyday wording. BM25 has an advantage on the first; a good embedding model may help more on the second. Keeping both candidate lists gives the reranker a better chance to see the right passage.
I avoid adding raw BM25 and cosine scores directly, because their scales mean different things. LumenTrail uses reciprocal rank fusion. Each passage earns a small score from its position in each list:
# A fragment of lumentrail.retrieval.reciprocal_rank_fusionfor position, passage_id in enumerate(ranking, start=1):scores[passage_id] = scores.get(passage_id, 0.0) + 1 / (60 + position)
This is a good default to test, not a promise of improvement. If the keyword index, dense model or chunks are poor, merging them can still return poor results. A cross encoder reranker can read a question and candidate passage together, which often helps ranking, but it adds latency and has its own blind spots. LumenTrail offers an optional general passage reranker; it is not trained to approve clinical claims.
With a small collection, I compare the query to every vector. That exact search gives me a reference result. As the collection grows, scanning every vector for every question may become too slow. An approximate nearest neighbour index searches a promising part of the collection instead. It saves time by accepting a chance of missing a neighbour.
| Index | What it does | When I would test it |
|---|---|---|
| Exact scan | Compares against every vector. | First baseline, small collections, or strict recall checks. |
| HNSW | Moves through a graph of nearby vectors. | A strong first approximate index when memory is available. |
| IVFFlat | Groups vectors and probes selected groups. | Larger collections where training and probe tuning make sense. |
| IVF with product quantisation | Searches groups using compressed vectors. | When index memory is a real constraint. |
| Disk oriented index | Keeps much of the search structure on fast storage. | Collections too large for a comfortable memory footprint. |
HNSW can offer a good speed and recall balance, but the graph takes memory and time to build. IVF needs enough data to form useful groups; probe too few and it misses results. Product quantisation saves space while changing the accuracy trade-off. The Faiss index guide is a useful map of these choices. pgvector supports exact search, HNSW and IVFFlat inside PostgreSQL, which can be a convenient fit when the documents and access metadata already live there.
Filters change the picture. A fast HNSW result over every document says little about a query restricted to one tenant or a narrow date range. Test recall and latency with the real filters, updates and concurrency you expect. If a filtered approximate search returns too few results, look at the index’s search settings and how filtering interacts with it before increasing k blindly.
I would choose around the data and the team that will operate it. An index benchmark alone misses permissions, updates, backups and the work of keeping source text in step with vectors.
| Starting point | Why it might fit | What to check before committing |
|---|---|---|
| PostgreSQL with pgvector | Documents and access rules already live in PostgreSQL. | Filtered search plans, index size and write load. |
| Elasticsearch or OpenSearch | The team already runs a search stack and needs strong text search alongside vectors. | How hybrid ranking, filters and operations behave in your version. |
| Azure AI Search | An Azure application needs managed text, vector and hybrid search. | Index design, identity, region and running cost. |
| Qdrant, Weaviate or Milvus | A dedicated vector service fits the collection and operating model. | Exact filter behaviour, hybrid options, backups and portability. |
| Pinecone or MongoDB Atlas Vector Search | A managed service fits the team’s existing platform. | Data location, export, access rules, hybrid features and price. |
| Faiss in an application | A local index or research baseline is enough. | How you will store metadata, enforce access and update the index. |
These products do not expose identical search or security features. For example, Elasticsearch, Azure AI Search, Qdrant and Weaviate each document a hybrid route, but their ranking and query controls differ. Prototype the same labelled questions and filters in the candidates you can actually maintain. A local SQLite index, like LumenTrail’s, is a useful place to learn the mechanics; it is not a multi-tenant search service.
I start with the questions people actually ask and the material they need to find. A compact model is attractive for a laptop or a Raspberry Pi, but a small English model may be a poor fit for a multilingual library. A model that does well on broad web passages may miss clinical abbreviations, negation or subtle differences between treatments.
LumenTrail uses BAAI/bge-small-en-v1.5 as an optional local model because it is light enough for a hands-on experiment. When the collection spans languages or longer passages, BGE-M3 is another candidate to evaluate, with higher compute and storage needs. I would also compare a suitable domain model if medical terminology dominates. The deciding result should come from reviewed questions in the intended language and domain, not a general leaderboard position.
Changing an embedding model means re-embedding the stored passages. Keep the model name and version with the index, and use the model’s recommended query formatting. Mixing vectors from different models in one search space can produce results that look plausible while being meaningless. LumenTrail rejects a query when its selected model does not match the one used to build the index.
A good result today can become a bad result when a page changes. I keep the original source ID, version or retrieval date, ownership and licence with each passage. I remove duplicates before splitting documents, preserve headings and cautions in the right chunk, and keep a route from every hit back to the source. For health material, I check who published it and when it was reviewed. A dataset mirror is a convenient download, not proof that its licence or content matches the original.
I also plan for change before filling the index. A corrected page needs a new vector and a retired old one. A deletion has to reach the keyword index, vector index, caches and any stored context. If access rules change, filtering must happen before private text reaches a model. I test those paths with real filters and record how long updates take. When I switch embedding models, I build a separate index and compare it on held-out questions before moving traffic.
Questions and retrieved passages may contain personal information even when the source collection is public. I collect only what the service needs, set retention rules for logs and memory, and keep private journal notes separate from shared documents. Retrieved pages can also carry malicious instructions; the application should treat them as untrusted content, whether the reader is an answer model, Codex or Claude Code.
A vector search finds nearby passages. It may struggle when a question asks how several people, documents or events connect. An entity graph can follow those relationships. Microsoft’s GraphRAG uses graph structure for entity focused search and broader questions about a collection. Building that structure costs time and can introduce extraction errors, so I compare it with simpler retrieval before adopting it. LumenTrail has a small, explicit entity expansion exercise to show the idea; it is not a full GraphRAG system.
Context is the evidence selected for one model request. More context is not always better. Irrelevant passages can distract the model, consume tokens and hide a useful warning. I keep the source ID with each passage and check whether every claim in the answer is supported by the context actually supplied.
Memory is different from a document library. A person’s dated note about a symptom or goal belongs to that person, can change over time, and should be deletable. It should not silently join a shared index. LumenTrail’s temporal journal uses synthetic notes only; it is a learning example, not a patient record system. The same boundary matters when Codex or Claude Code reads retrieved material: source text is data to inspect, never an instruction with authority over the agent.
The pieces fit together without requiring every project to use every piece:
public source → clean and chunk → BM25 + vectors → merge → rerank → cited context↓person's dated notes → owner, consent and validity check ─────→ optional answer model↓answer with source IDs
Evaluation watches the whole route. It checks whether ingestion preserved the right text, whether search found the labelled passages, whether memory stayed with its owner, whether the answer used its citations correctly, and how long each stage took. It is not simply a score attached after the answer. The evaluation guide walks through those separate checks.
I would add memory only when the task benefits from it and the user expects it. pgvector is a PostgreSQL extension for vector search. Graphiti is a framework for temporal facts and relationships, with its own graph storage needs. They are different choices, not two names for the same memory layer. Start with an explicit, user-scoped note and a retrieval baseline. Move to extracted facts or a graph when multi-session tests show that simpler state is insufficient.
LumenTrail implements the chunking exercise, hybrid retrieval, optional reranking, scoped synthetic notes and retrieval metrics. It does not generate answers. That keeps source inspection possible before adding a model that can make fluent but unsupported claims.
The original MedQuAD collection contains 47,457 question and answer pairs from 12 NIH sites. Three MedlinePlus subsets omit answer text for copyright reasons, and the remaining material still needs a freshness check before anyone treats it as health guidance. LumenTrail can fetch a small sample from the original GitHub source. It can also adapt a CSV from a Kaggle MedQuAD mirror, provided you check the mirror’s fields, licence and provenance.
HealthSearchQA supplies consumer worded questions. It does not supply the gold passages needed to score retrieval against your own corpus. That makes it useful for finding natural question styles, but a reviewer still has to say which passages are relevant.
lumentrail initlumentrail search "What is AF?" --method keywordlumentrail search "What is AF?" --method hybridlumentrail eval --method hybrid --limit 2
On the tiny synthetic set, these commands check that the code runs and the known passage appears. They do not measure medical safety. For a more useful experiment, import public data, review a holdout set of questions, compare keyword and real dense retrieval, then inspect every miss. I walk through that process in How to Evaluate RAG Without Fooling Yourself.
When I build a real service, I also test source deletion, access changes, stale pages, exact names, no-answer questions, answer citations, latency and cost. Those checks tell me far more than a screenshot of a nearest-neighbour result.
Legal Stuff
