18 March 2026
Retrieval quality is a citation problem before it is a model problem
When a retrieval system fails, the first instinct is to swap the model. In the engagements we have run this year, the miss was more often a citation miss: the passage was in the corpus, the answer sounded right, and the file it pointed at was the wrong annex, the wrong year, or a slide that summarised a clause the contract had already replaced. Embeddings, the numerical summaries that let similar passages sit near each other, will happily retrieve a near neighbour. A near neighbour can still be the wrong clause for a partner to sign.
We now score three things on every evaluation question. Did the system refuse when the files had no answer. Did it cite a page a person agrees is the right source. Did the answer stay inside that page, or did it drift into fluent filler. The model matters for the third. The first two are design: chunk size, metadata, permissions, and the stubborn requirement that a reply without a citation is a miss even when a reviewer likes the prose.
A practical consequence is that we spend week two on the files. The prompt comes after the corpus can be trusted. Duplicate PDFs, scanned schedules without page numbers, and matter folders that mix drafts with executed copies will wreck a beautiful model. Cleaning is part of the retrieval product, with a size and an owner. If a folder cannot be trusted, we leave it out of the index and write that down, rather than hoping the embedding space will sort it out.