Chunking algorithms divide a large input into smaller, usable units called chunks. In software engineering, this may mean splitting files or workloads for processing. In natural language processing (NLP) and retrieval-augmented generation (RAG), chunking usually means breaking documents into passages that can be embedded, indexed and retrieved by an AI system.
The quality of your chunks directly affects search relevance, embedding accuracy, context-window usage, latency and answer quality. Chunks that are too large may contain unrelated information; chunks that are too small may lose essential context. The right algorithm depends on the data structure, downstream model, query patterns and operational constraints.
What Are Chunking Algorithms?
A chunking algorithm applies a repeatable rule to divide content into smaller segments while preserving as much useful meaning as possible. A chunk can be defined by:
- A fixed number of characters, tokens or bytes
- Natural boundaries such as paragraphs, sentences or headings
- Semantic similarity between adjacent text segments
- Structural units such as HTML elements, pages, tables or code blocks
- Processing constraints such as batch size, memory or file limits
Chunking is not merely a preprocessing task. For AI applications, it is part of information architecture. The algorithm determines what an embedding represents, which passage is returned for a query and how much evidence an LLM receives in its prompt.
A useful chunking pipeline typically includes:
1. Document loading — Read PDFs, web pages, office files, databases or APIs.
2. Cleaning and normalization — Remove boilerplate, repair encoding and preserve meaningful structure.
3. Segmentation — Apply one or more chunking algorithms.
4. Metadata enrichment — Add title, section, page, source, timestamp and access-control fields.
5. Embedding or indexing — Store vectors, keywords or both.
6. Evaluation — Measure retrieval and answer quality using representative queries.
Why Chunking Matters in AI and RAG Systems
A retrieval system can only return the units it has indexed. If a key definition is split from its conditions, a retrieved passage may be technically relevant but incomplete. Conversely, if every chunk contains several unrelated topics, vector similarity becomes less precise.
Good chunking can improve:
- Recall: The correct evidence is available to the retriever.
- Precision: Retrieved passages focus on the user’s question.
- Context utilization: Prompt tokens contain useful information rather than repetition.
- Citation quality: Answers can point to specific sections or pages.
- Latency and cost: Smaller, targeted contexts reduce model input tokens.
- Update efficiency: Only changed chunks need to be re-embedded.
For Indian businesses, chunking is especially relevant when processing multilingual documents, scanned government forms, GST or compliance records, internal policies and customer-support content. English-only assumptions can fail on Hindi, Tamil, Bengali and code-mixed material, where tokenization and sentence boundaries differ.
Fixed-Size Chunking
Fixed-size chunking divides text into units of a predetermined length. The length may be measured in characters, words or model tokens. A common implementation uses a chunk size plus overlap.
For example, a system might create chunks of 500 tokens with a 50-token overlap. The overlap repeats the boundary context so that a sentence or concept split between two chunks remains visible in at least one retrieved unit.
Character-based chunking
Character-based chunking is easy to implement and fast. It is useful for raw text, logs and simple ingestion pipelines, but character counts do not correspond consistently to model tokens or meaning. A 2,000-character chunk may contain very different amounts of information depending on language, formatting and punctuation.
Token-based chunking
Token-based chunking measures text according to the tokenizer used by the target embedding or language model. This provides more reliable control over context windows and API limits. However, tokenization varies by model, and Indic scripts, emojis, URLs and source code can produce unexpected token counts.
Example in Python
def fixed_chunks(tokens, size=500, overlap=50):
if overlap >= size:
raise ValueError("overlap must be smaller than size")
step = size - overlap
return [tokens[i:i + size] for i in range(0, len(tokens), step)]Fixed-size chunking is a strong baseline because it is predictable and inexpensive. Its main weakness is that it can cut through headings, tables, lists and explanations. Use it when document structure is poor, speed matters or you need a benchmark for more advanced methods.
Sentence-Based Chunking
Sentence-based chunking splits content at sentence boundaries and then groups sentences until a target length is reached. This avoids many awkward cuts produced by character-based methods.
A typical process is:
1. Detect sentence boundaries.
2. Add sentences to the current chunk.
3. Stop when the token or character budget is reached.
4. Optionally retain one or more previous sentences as overlap.
Sentence detection must account for abbreviations, decimal numbers, titles, legal references and multilingual punctuation. A naive split on full stops can incorrectly separate “Dr.”, “No.” or decimal values.
Sentence-based chunks work well for news, FAQs, policy documents and prose. They are less effective for spreadsheets, code, slide decks and documents where meaning depends heavily on layout.
Paragraph and Recursive Chunking
Recursive chunking attempts several separators in order of importance. For example, it may first split by headings, then paragraphs, then sentences and finally words or tokens if a segment is still too large.
A common hierarchy is:
Heading
→ paragraph
→ sentence
→ word or tokenThe algorithm preserves larger semantic units whenever they fit within the size limit. If a section exceeds the limit, it recursively applies a finer separator.
Pseudocode
def recursive_split(text, separators, max_chars):
if len(text) <= max_chars:
return [text]
for separator in separators:
parts = text.split(separator)
if len(parts) > 1:
chunks, current = [], ""
for part in parts:
candidate = current + separator + part if current else part
if len(candidate) <= max_chars:
current = candidate
else:
if current:
chunks.append(current.strip())
current = part
if current:
chunks.append(current.strip())
return chunks
return [text[i:i + max_chars] for i in range(0, len(text), max_chars)]Recursive chunking is widely used because it provides a practical balance between simplicity and structural awareness. It should still be adapted to the source format. Markdown headings, HTML tags, PDF layout and code syntax need different separator hierarchies.
Semantic Chunking
Semantic chunking creates boundaries when the meaning of the text changes significantly. Instead of relying only on punctuation or length, it compares adjacent sentences or groups using embeddings.
One approach is:
1. Split the document into sentences.
2. Create an embedding for each sentence or sentence group.
3. Calculate similarity between adjacent groups.
4. Mark a boundary when similarity falls below a threshold.
5. Enforce minimum and maximum chunk sizes.
If a document moves from product features to pricing, a semantic algorithm may detect that transition even without a new heading.
Semantic chunking can improve retrieval for heterogeneous documents, but it has costs:
- Sentence-level embedding calls increase processing time and expense.
- Thresholds are domain-dependent.
- Similar sentences may still require different metadata or permissions.
- A purely semantic boundary can create chunks that are too small or too large.
In production, combine semantic decisions with hard size limits. A semantic chunk should never exceed the embedding model’s safe input capacity or the context budget of the generation model.
Structure-Aware Chunking
Structure-aware algorithms use the document’s native organization. This is often the best choice when structure is reliable.
Examples include:
- HTML: Split by article, headings, sections, lists and tables while removing navigation and advertisements.
- Markdown: Use heading levels, code fences, lists and block quotes.
- PDF: Preserve page numbers, headings, paragraphs, table regions and reading order.
- Office documents: Retain paragraphs, styles, headers, footers and tables.
- Code: Split by modules, classes, functions and documentation comments.
- JSON or XML: Chunk records or logical nodes rather than raw characters.
- Transcripts: Use speaker turns, timestamps and topic changes.
For technical documentation, a useful chunk often includes the section heading in every child chunk. This supplies context without repeating the entire document. For tables, store a compact title and header representation with each row group so that a retrieved row remains interpretable.
Overlap: How Much Is Enough?
Overlap protects against boundary loss, but excessive overlap increases index size, duplicate retrievals and prompt costs. The correct amount depends on how self-contained your content is.
Typical starting points are:
- 5–15% overlap for well-structured prose
- One or two sentences for FAQ and policy content
- A small number of lines for code
- Row or header repetition for tables
- Timestamp or speaker overlap for transcripts
Overlap should not be used to compensate for poor chunk boundaries. If a definition and its exceptions are routinely separated, improve the structural or semantic algorithm first. Then use overlap as a safety measure.
Choosing Chunk Size
There is no universal best chunk size. Start with the retrieval task rather than a popular number.
Consider:
- Embedding model limits: Stay below the model’s input token limit, with safety margin.
- Answer complexity: Multi-step questions may require larger evidence units.
- Document density: A legal clause and a marketing paragraph contain different information per token.
- Query length: Short queries often benefit from focused chunks.
- Reranking: A reranker can handle a larger candidate pool but still benefits from coherent passages.
- LLM context budget: Reserve tokens for instructions, conversation history and the generated answer.
A sensible testing range might compare 200, 400 and 800 tokens, with controlled overlap. Do not evaluate only retrieval similarity. Measure whether the returned chunks contain sufficient evidence to answer real questions.
Hybrid Chunking Strategies
Production systems frequently combine algorithms. A robust hybrid pipeline might:
1. Parse document structure.
2. Split at headings and paragraphs.
3. Recursively divide oversized sections.
4. Apply sentence-aware overlap.
5. Use semantic boundaries only within very long sections.
6. Attach parent-section metadata to every child chunk.
Another effective pattern is parent-child retrieval. Index small child chunks for precise matching, but return the child plus its parent section or neighboring chunks to the language model. This separates retrieval granularity from generation context.
For multilingual Indian datasets, consider language-aware sentence segmentation and evaluate each major language separately. A chunk size that works for English may not work for Hindi or Kannada because tokenization, morphology and script affect both length and semantic density.
Common Chunking Mistakes
Splitting without metadata
A chunk without its title, page, source or date may be impossible to interpret. Store provenance and access-control metadata with every chunk.
Ignoring document layout
PDF extraction can scramble columns, headers and footnotes. Validate extracted text before choosing a chunking algorithm.
Using one configuration for every source
Support tickets, contracts, code and manuals have different logical units. Configure chunking by content type.
Overlapping too aggressively
Large overlap creates duplicates and inflates vector-store costs. Measure its effect instead of assuming more is better.
Indexing boilerplate
Repeated navigation, disclaimers and email signatures can dominate retrieval. Remove or downweight boilerplate during ingestion.
Failing to preserve access controls
In enterprise RAG, chunk retrieval must respect document permissions. Metadata filters should be applied before or during retrieval, not after an LLM has seen unauthorized content.
Evaluating Chunking Algorithms
Evaluation should use a representative query set covering factual lookup, multi-hop questions, terminology variations, languages and difficult edge cases.
Useful metrics include:
- Recall@k: Whether a relevant chunk appears in the top k results.
- Precision@k: How many returned chunks are relevant.
- MRR or nDCG: Ranking quality across result positions.
- Answer faithfulness: Whether the generated answer is supported by retrieved evidence.
- Context precision: How much supplied context is actually useful.
- Index cost: Number of chunks, storage and embedding volume.
- Latency: Ingestion and query-time performance.
Keep the retriever, embedding model, query set and ranking settings constant while comparing chunking strategies. Otherwise, you cannot tell whether an improvement came from chunking or another change. Log chunk IDs, source locations and retrieval scores so errors can be inspected manually.
Practical Implementation Checklist
Before deploying a chunking pipeline, verify that:
- The text extractor preserves reading order and Unicode correctly.
- Chunk lengths are measured with the target model tokenizer where possible.
- Every chunk has stable IDs and source references.
- Headings and parent context are retained.
- Tables, code and lists have specialized handling.
- Overlap is bounded and tested.
- Metadata filters enforce tenant and permission boundaries.
- Re-ingestion produces deterministic chunk IDs when content is unchanged.
- Evaluation includes multilingual and adversarial examples.
- Changes to chunking configuration trigger controlled re-indexing.
Frequently Asked Questions
What is the best chunking algorithm for RAG?
Recursive or structure-aware chunking is a strong default for most RAG systems. Add token limits, modest overlap and parent metadata, then compare alternatives using a real evaluation set.
Are larger chunks better for AI search?
Not necessarily. Larger chunks preserve context but can reduce retrieval precision and consume more prompt tokens. The optimal size depends on document density, query type and model behavior.
Should chunks overlap?
Usually, yes—but modestly. Sentence or section-aware overlap is preferable to a large fixed overlap. Test whether it improves recall enough to justify extra storage and duplicated context.
Can chunking algorithms work with PDFs?
Yes, but reliable PDF extraction comes first. Preserve page numbers, reading order, headings and table boundaries; otherwise, even an advanced algorithm will chunk corrupted text.
How do I chunk multilingual documents?
Use language-aware sentence segmentation, tokenize with the target model and evaluate each language separately. Preserve original text and, where appropriate, language metadata for filtering and analysis.
Apply for AI Grants India
Building an AI product that uses retrieval, NLP or intelligent data processing? Apply to AI Grants India for support, visibility and opportunities designed for Indian AI founders.