A RAG system cannot compare a question directly with every passage by asking a language model to read the entire collection each time. In a common vector-based design, it creates an embedding for each stored chunk: a fixed-length list of numbers produced by an embedding model. When a question arrives, the system creates a compatible embedding for the query and searches for nearby chunk vectors.
Earlier RAG Explainers covered How RAG Works: From Document Indexing to Generated Answers and Why RAG Systems Split Documents into Chunks.
The next step in the document-preparation path is to understand how embeddings support retrieval. The focus here is the mechanism, without treating vector search as the only retrieval method or claiming that a high similarity score guarantees a useful answer.
An embedding is a numerical representation
An embedding is a vector: an ordered list of numbers created from an input such as a sentence, paragraph, image, or query. For this article, the inputs are text. The model maps each text into a space where related inputs are intended to have representations that can be compared.
The individual numbers are not usually readable labels such as “laptop,” “deadline,” or “policy.” Meaning is distributed across the vector, and the representation depends on the model that produced it. A diagram may draw vectors as points on a flat page, but production embeddings commonly have many dimensions. The drawing is a mental model, not a literal map of the stored coordinates.
During indexing, the application keeps the readable chunk text and its source metadata alongside the chunk’s vector representation. The vector supports comparison; the original text is what the application can later supply to the language model.

How the query is compared
The stored chunk vectors and the query vector must come from a compatible encoding setup. In a basic system, the same embedding model processes both. Some retrieval models use different query and document instructions or routes, but their outputs are trained to occupy a comparable space.
The search system then applies a measure such as cosine similarity, dot product, or Euclidean distance. The exact convention matters: a larger similarity value may mean closer, while a smaller distance may mean closer. The system uses that measure to order candidates and usually returns a limited number of the nearest chunks.
This comparison happens between vectors, but the output of retrieval should still identify the corresponding chunk text and source. The model answering the question does not need a row of floating-point numbers; it needs the selected evidence in readable form.
Example: matching different wording
Consider a fictional employee handbook. One chunk says, “Departing employees must return company laptops within five business days of their final working day.” A reader asks, “When do I need to give back my work computer?” The question and passage share little exact phrasing beyond the general subject.
An embedding model may place their vectors near each other because “give back” relates to “return,” “work computer” relates to “company laptop,” and the question asks for a time condition that the passage contains. A second chunk about ordering laptop accessories shares the word “laptop” but does not address the deadline. Vector similarity can therefore surface a passage even when simple word overlap is limited.

This remains model-dependent behavior. The example illustrates the purpose of semantic retrieval; it does not claim that every embedding model will rank these passages in the same order.
Why the nearest result can still be insufficient
Now consider a fictional access guide. A user asks, “Can I keep using my old phone after moving authentication to a new one?” The collection contains one chunk about enrolling a replacement phone and another about reporting a lost device. Both are related to phones and authentication, but neither states whether the old phone remains usable.
A similarity search can still rank one of those chunks first because “nearest” is defined relative to the available candidates. If the collection lacks a passage that answers the question, the top result may be related but incomplete. Returning several neighbors may expose more context, yet it does not create a missing policy.
This is why retrieval and answer generation remain separate stages. Similarity search chooses candidates according to a representation and comparison rule. The application and language model must still work with the evidence that was actually returned.
What a similarity score does not mean
A similarity score describes a relationship between vectors under a particular model and comparison method. It is not a probability that the passage is correct, current, complete, or sufficient for the question. A cosine similarity of 0.82, for example, should not be read as “82 percent likely to answer correctly.”
Raw values also do not have a universal meaning across embedding models, distance functions, or index configurations. Even within one system, a score is most naturally used to order candidates or apply a rule that the application has separately evaluated. The value alone does not explain why two texts were placed near each other.
Other retrieval choices still matter: which chunks exist, which candidates are eligible, how many are returned, and whether another stage filters or reranks them. This article stops at the vector comparison itself; those later mechanisms deserve their own explanations.
During indexing, each readable chunk is paired with a vector representation. At query time, the question is encoded into a compatible vector. Similarity search compares that query vector with stored vectors, orders nearby candidates, and returns the associated text and source information.
Embeddings make a kind of semantic comparison possible; they do not determine whether the selected passage is true or sufficient. The next article will compare dense, sparse, and hybrid retrieval mechanically, including what information each method uses when producing candidates.


Practical add from running pgvector in production: similarity scores are only meaningful relative to your own data, so tune thresholds empirically instead of trusting defaults. What worked for me was logging score distributions on real queries, then picking cutoffs from that, not from documentation. And if you haven't tried hybrid retrieval yet, it's worth it: dense embeddings miss exact-term matches that sparse retrieval catches easily, and combining them fixed a whole class of "the answer was right there" failures for me. The theory is clean; the tuning is where the wins are.