A RAG system usually searches pieces of documents rather than entire files. A long handbook may contain dozens of unrelated policies; retrieving the whole handbook for a question about one policy would bring along a great deal of unrelated text. Dividing it into chunks gives the system smaller units to index and return.
That division is a design choice, not merely a formatting step. It decides which words are stored together, what a search result can contain, and what context the model will see if that result is selected. The aim here is to understand the mechanism, not to prescribe a universal chunk size.
What a chunk actually contains
A chunk is a retrievable unit derived from a source. For text, it is often a passage from a document, together with an identifier and metadata such as the source title or section. A system might split at paragraphs or headings, or use a length limit measured in characters or tokens. Different implementations combine these signals in different ways.
Consider a fictional employee handbook with sections on equipment, travel, and leave. The equipment section can become one or more chunks. In a vector-based system, an embedding of each chunk can help find related material, while the readable chunk text remains available for the answer. Other search methods may represent and retrieve the text differently; chunking is not exclusive to vectors.

The source document still matters. A chunk’s text may be short, but a source identifier and section label can show where it came from. Those labels must be retained accurately during preparation; a model cannot recover missing provenance simply because it received the passage.
Why the boundary matters
A chunk boundary controls which nearby words travel together into search and, later, into the model’s context. If a sentence is cut between a subject and its condition, a single returned piece may not carry the complete fact. If a chunk spans several unrelated sections, it can carry material the question did not ask for. These are possible effects of a boundary, not guaranteed failures of any particular splitter.
Take the handbook statement: “Departing employees must return laptops within five business days of their final working day.” A cut after “within” puts the deadline in the next chunk. If retrieval returns only the first piece, the answer lacks the deadline. Keeping the sentence together makes that specific fact available in one piece. This is an illustration, not a measured comparison of chunking methods.

A chunker can use document structure when that structure is available: headings, paragraphs, lists, or other boundaries. In a messy PDF, extraction may lose those signals, so the usable boundaries depend partly on parsing. A clean source hierarchy can help, but it cannot guarantee that every useful fact fits in one chunk.
Overlap and source context
Some chunkers repeat a small stretch of text in adjacent chunks. That is overlap: the end of one chunk appears again at the start of the next. Repetition can keep a phrase near a boundary visible in either neighboring piece, but it also creates duplicate indexed text. Overlap changes the available units; it does not ensure that the retriever will select the right one.

Metadata provides a different kind of context. A passage headed “Equipment returns” may be easier to interpret when its heading and source title accompany it. A system may store those as fields, include them in the searchable text, attach them to the model input, or use some combination. Those choices affect search and presentation differently, so “the chunk includes metadata” should not be taken to mean the model always sees every field.
The original document remains the source of truth. A retrieved chunk is a selected excerpt with whatever surrounding material the application has deliberately retained or fetched. It should not be mistaken for the whole document.
A second example: a procedure with a prerequisite
Imagine a fictional operations guide with a section titled “Rotate a service key.” Its first paragraph says to create a replacement key and verify that a test request succeeds. The next paragraph begins, “After that check, revoke the old key.” A question asks, “When can the old key be revoked?”
If those paragraphs are stored separately and retrieval returns only the second, “that check” has no explanation in the selected text. A chunk containing both paragraphs, or a retrieval step that also fetches the preceding passage, can supply the missing prerequisite. The answer depends on what was actually retrieved and passed to the model; chunking alone does not force the model to honor the sequence.

This example is synthetic. It shows why a boundary and the amount of adjacent context can change the evidence available for one question. It says nothing about the best general settings for operational guides.
What chunking does not solve
Chunking cannot repair text that was never extracted. If a parser drops a table column, cuts off a footnote, or confuses reading order, the index may contain an incomplete or misleading passage. A well-placed boundary around that passage does not restore what is absent.
Nor does a chunk guarantee an answer. Search may select the wrong piece, omit a neighboring condition, or return several pieces that disagree because their sources describe different versions. The application still has to decide what to retrieve, what to pass to the model, and how to present the source. These are separate stages of the system.
Chunking is therefore one control over the evidence available at retrieval time. It is most useful to think of it as defining the granularity of a search result and the local context that result carries.
The source document is prepared; the chunker chooses retrievable units; the index stores those units and their representations; the retriever selects some for a question; and the application gives selected text to the model. A boundary determines which words can arrive together as one result, while metadata can preserve a link to the source.
There is no single boundary that is right for every document and question. The practical first step is to see which units a system created and what text a real retrieval result contains. The following article will explain embeddings and similarity search, another part of how systems find those units.

