BETA HEALTH / CASE STUDY

Making Medical Records Easier to Understand with RAG

How I built the retrieval layer behind Beta Health’s medical-record assistant, and the engineering decisions that shaped it.

Making Medical Records Easier to Understand with RAG

01

THE PRODUCT

Beta Health is a patient-facing healthcare application that gives patients access to and a better understanding of their health information. The medical-record assistant is one capability within that product, helping patients ask questions about the records they already have access to.

02

THE PROBLEM

Beta Health already gave patients access to their medical records. The problem was what happened after that. A patient could open a visit and navigate summaries, notes, lab orders, tests, and nested results, but the structure was not always how a patient thought about their records. A patient was more likely to ask: “What were my blood test results?” or “What changed between these two visits?”

03

1. SCOPE BEFORE SEARCH

I did not want every question searching the patient’s entire medical history. Repeated investigations across different visits can all be semantically strong matches. I moved the first decision to the patient: before asking questions, the patient selects the visits they want the conversation to use. Retrieval then searches inside that selected context.

04

2. ON-DEMAND INDEXING

A visit is indexed when the patient first wants to use it. Indexing every historical visit would create embeddings that might never be used; embedding on every question would make repeated questions slow and expensive. I track NOT_INDEXED, QUEUED, INDEXING, READY, FAILED, and STALE states. The vector index grows according to what patients actually use.

05

3. PARENT AND CHILD CHUNKING

My biggest retrieval problem was not embeddings. It was deciding what exactly to embed. A lab order is not a flat document; it can contain multiple tests and nested results. I separated clinical parents from search children. Smaller children improve search precision, while the matching child resolves back to a larger clinical parent for generation context. This keeps a value connected to the test, unit, order, and result that give it meaning.

06

4. THE QUERY PATH

The vector query is restricted to the selected visits. Matching children are resolved back to their clinical parents before generation. Gemini receives the question and retrieved evidence rather than the patient’s entire history.

07

5. PRESERVING THE RECORD

Generation instructions explicitly preserve values, units, dates, medication names, negations, uncertainty, and conflicting records. “Suspected malaria” must not become “You were diagnosed with malaria.” If two selected records disagree, the assistant should show the disagreement rather than invent a conclusion. Example: “Your malaria result was recorded as positive on 12 August [S1] and negative on 18 August [S2].”

08

6. CITATIONS AS PART OF RETRIEVAL

Citations are not added after generation. Each retrieved clinical parent already has a source identity such as [S1]. The API separately returns the visit, section, record ID, and recorded-at metadata so a citation can take a patient back to something recognizable in the original record.

09

7. REINDEXING

Persisted indexes eventually need rebuilding when source records change, chunking changes, or the embedding model or vector dimensions change. Replacement happens transactionally so retrieval never sees a mixture of an old and new index. Either the complete replacement becomes available or the previous READY state remains.

10

8. WHAT IS NOT SOLVED YET

The biggest remaining area is evaluation. I need a proper evaluation set with expected records for repeated results, similar test names, unanswered questions, negations, uncertain diagnoses, conflicting notes, date-sensitive questions, and questions that require multiple records. Citation validation also needs to check that a source supports a claim, not only that the source label exists.

11

WHAT I TOOK AWAY

The consequential decisions happened around the LLM: which visits to scope, when to index, what counts as a clinical unit, how to retrieve, how to expand context, what the model is allowed to claim, and how to preserve provenance. The LLM is only one part of the system; most of the engineering work is controlling what reaches it and what comes back out.