Text Splitter: The Complete Guide (2026)
The units you can split by, why sentence boundaries are harder than they look, what the published guidance on chunking for retrieval recommends, how tokens differ from characters, and how to keep parts in order once they travel separately.
- Why splitting text comes up everywhere
- Seven units of splitting, and when each fits
- Sentence boundaries are harder than they look
- Chunking for retrieval and embeddings
- Fixed-size, recursive and semantic chunking
- Overlap: what it does and how much to use
- Splitting by tokens: why characters are not enough
- Context windows are not the limit you think
- Platform limits: SMS, posts and fields
- Labels, order and putting the parts back together
- A repeatable workflow
- Related tools for the same job
Splitting text is the quiet prerequisite of half the things people do with language today. A prompt has to fit a context window, a document has to be chunked before it can be embedded and searched, a post has to fit a platform’s cap, a message has to fit an SMS, a transcript has to be reviewed by three people at once. Each of those is the same problem: cut a long text into parts that each fit a limit, without cutting where it hurts. This guide covers the units you can split by, why sentence boundaries are harder than they look, what the published guidance on chunking for retrieval actually recommends, how tokens differ from characters, and how to keep the parts in order once they travel separately.
The Text Splitter does all of this in your browser. The guide explains the decisions behind each of its options so you can choose them on purpose.
Why splitting text comes up everywhere
Limits are everywhere, and they are measured in different units. Language models measure their context in tokens. Embedding models have a maximum input, also in tokens. Text fields, SMS and social platforms count characters. Editors, translators and reviewers think in words, sentences and paragraphs. Logs and spreadsheets think in lines. A splitter is useful precisely because it lets you cut by the unit the destination counts, rather than by the one that happens to be easy to measure.
The second reason splitting matters is quality, not just fit. A retrieval system that answers questions from your documents can only return the chunks it indexed; if a chunk ends mid-sentence, the fact on the boundary is in neither half. Pinecone’s guide to chunking strategies puts it directly: if a chunk contains sentences that are not useful without context, it may not be surfaced when querying. Splitting well is a way of deciding what a future search will be able to find.
Seven units of splitting, and when each fits
| Unit | What the size means | Use it when |
|---|---|---|
| Characters | Maximum characters per part | The destination counts characters: form fields, SMS, posts, listings |
| Words | Maximum words per part | People will read or translate the parts; word counts are how they plan work |
| Tokens | Maximum tokens per part in the model’s own encoding | The parts go to a language model or an embedding model |
| Sentences | Sentences per part | You want the smallest unit that still reads as a complete thought |
| Paragraphs | Paragraphs per part | The author already grouped ideas; one paragraph per chunk is a common retrieval unit |
| Lines | Lines per part | The text is already one item per line: logs, lists, CSV rows, code |
| Equal parts | The number of parts you want | You are dividing a job among people or sessions and only the count matters |
Two of these deserve a note. Paragraph splitting respects a structure a human already chose, which is why it is often the best first attempt for well-edited documents. Equal parts is the odd one out: it starts from the number of parts and derives the size, then still ends each part on a sentence or word boundary so that no part begins mid-thought.
Sentence boundaries are harder than they look
Everyone knows a sentence ends with a period, a question mark or an exclamation mark. The trouble is that the period also ends abbreviations (“Dr.”, “U.S.”, “etc.”) and sits inside numbers (“3.14”), and quotation marks and closing parentheses can follow the terminal punctuation. The Unicode standard for text segmentation, UAX #29, says so plainly: the period is used ambiguously, sometimes to end a sentence, sometimes for abbreviations, sometimes for numbers. Its default sentence-boundary rules handle the common cases, and it recommends language-specific tailoring for the rest.
Browsers now ship that machinery. The Intl.Segmenter API performs locale-sensitive segmentation into graphemes, words or sentences, so a splitter running in the browser can ask the platform where sentences end instead of guessing with a regular expression. The Text Splitter uses it when available, with a rule-based fallback for older browsers, and packs whole sentences into each part. When a single sentence is longer than the part you asked for, it has no choice but to cut inside it, and it tells you so rather than silently doing it.
Chunking for retrieval and embeddings
Retrieval-augmented generation, or RAG, indexes pieces of your documents as vectors and pulls the closest pieces into a prompt when a question arrives. The pieces are chunks, and the way you make them decides what the system can answer. Two constraints frame the choice. First, embedding models cap their input: OpenAI’s embeddings guide lists a maximum input of 8,192 tokens for its text-embedding-3 models. Second, a chunk should carry one idea with enough context to stand alone, because that is what the query is compared against.
Microsoft’s documentation on chunking large documents for vector search gives a concrete starting point: a fixed size sufficient for semantically meaningful paragraphs, for example 200 words or 600 characters, with some overlap, for example 10 to 15 percent of the content. It also notes that chunking is how you stay under the maximum token input of chat and embedding models. Those are defaults to test, not laws; a table of specifications wants smaller chunks than a narrative report, and questions that need cross-references want larger ones.
Fixed-size, recursive and semantic chunking
The chunking literature groups methods into a few families. Fixed-size chunking, as Pinecone describes it, means deciding a number of tokens per chunk and breaking the document into pieces of that size, optionally with overlap; it is the most common approach and the cheapest to compute. Recursive character splitting, implemented in LangChain’s RecursiveCharacterTextSplitter, tries a list of separators in order, typically paragraph breaks first, then line breaks, then spaces, so that chunks follow the document’s own structure as far as the size allows. Semantic chunking uses a model to decide where topics change, which is more expensive and cannot run without a model. LangChain’s text-splitter documentation is the reference for the software families; the Text Splitter implements the first two ideas in the browser (fixed size in any unit, with sentence- and word-aware boundaries) and deliberately not the third.
Overlap: what it does and how much to use
Overlap repeats the end of one chunk at the beginning of the next. Its purpose is to make sure that a sentence or fact sitting on a boundary appears whole in at least one chunk, so a search can still find it. The cost is redundancy: more chunks to store, more tokens to embed, and duplicated passages in results. Microsoft’s 10 to 15 percent figure is a reasonable starting range for retrieval; for text a person will read in order, overlap is simply noise and should be zero.
How the overlap is taken matters as much as how big it is. A raw slice of the last N characters usually starts mid-word. The Text Splitter carries back whole units instead, complete words or complete sentences depending on your settings, up to the overlap budget, so the repeated tail always reads as text and the boundary fact is preserved intact.
Splitting by tokens: why characters are not enough
A language model never sees your characters. It sees tokens, the units its tokenizer produces, and their relationship to characters is only a rule of thumb: in English prose one token is roughly four characters, but code, numbers, other languages and unusual words shift the ratio, sometimes by a lot. Our own tokenizer comparison measured how much the same text varies across encodings. A part sized in characters is therefore a guess when the limit is in tokens, and a guess that fails in the wrong direction gets truncated.
The Text Splitter counts with the o200k encoding, the one used by GPT-5, GPT-4.1 and GPT-4o, with the same vendored tokenizer as the Token Counter, so a part sized at 500 tokens is exactly 500 tokens for those models. It packs whole sentences or words up to the budget and then measures each finished part exactly. Claude and Gemini use their own tokenizers, which are not available as browser libraries; for them the count is a close estimate, and the right move is to size parts a little below the limit.
Context windows are not the limit you think
Context windows have grown enormously. Google’s documentation says many Gemini models come with context windows of 1 million or more tokens; Anthropic’s documentation describes 200,000-token windows for Claude Sonnet 4.5 and other models, and up to 1 million depending on the model. It is tempting to conclude that splitting is obsolete. It is not, for three reasons. Cost: every token in the window is paid for on every call. Quality: retrieval systems still return chunks, and a chunk that is a whole document is a poor search result. Attention: a model given one relevant paragraph usually does better than a model given a hundred pages and told to find it. Large windows change how much you can afford to include; they do not remove the need to choose.
Platform limits: SMS, posts and fields
Some limits are old and hard. A single SMS carries a 140-byte body, which holds 160 characters of the GSM 7-bit alphabet defined in 3GPP TS 23.038, or 70 characters when the message includes anything outside that alphabet and must be sent as UCS-2; a single accented letter or emoji can therefore halve the room and double the number of messages. X counts posts up to 280 characters and, per its counting-characters documentation, weights some characters, such as many CJK characters and emoji, as two. Meta descriptions, app-store fields, ad headlines and form inputs each have their own caps, and most of them count characters, not words. Character mode with whole words kept is the right tool for all of them; when a platform’s counting rules are unusual, check the largest part with the Character Counter.
Labels, order and putting the parts back together
Parts that travel separately need three things: a label that says which part this is and how many there are, a stable order, and a join rule that reverses the split. A label such as [2/5] costs six characters and prevents most confusion; the splitter adds it as a prefix so it is visible in any destination. Order is preserved by copying all parts at once or downloading them as a single file, in which the parts are separated by a blank line. To reassemble, remove the labels and join the parts with that same separator; if you used overlap, the repeated tails have to be removed as well, which is one more reason to use overlap only for retrieval and not for text a person will reassemble.
A repeatable workflow
- Name the limit and its unit. Tokens for a model, characters for a field, paragraphs for a document a person edited. Set the size a little below the real limit.
- Paste the text into the Text Splitter and pick the matching mode. Keep sentences and words whole unless the destination truly counts raw characters.
- Decide on overlap. Ten to fifteen percent for retrieval chunks, zero for anything a person reads in order.
- Look at the largest part. If one part is far bigger than the rest, a long sentence or paragraph forced it; lower the size or switch on word-level splitting.
- Keep the labels on when parts are sent separately, and copy all or download so the order is preserved.
- Count the destination’s way. Verify a part with the Token Counter or the Character Counter when the limit is strict.
Related tools for the same job
The splitter cuts; other tools measure and prepare. The Token Counter prices a text across models and checks it against a context window. The Sentence Counter and Word Counter show what you are about to split. The AI Text Cleaner removes Markdown and hidden characters from an assistant’s answer before you chunk it, and the Text Summarizer shortens a source when fitting it whole is the better choice.
Sources and further reading
- Microsoft Learn: Chunk large documents for vector search solutions (fixed size, overlap, embedding input limits)
- OpenAI: Embeddings guide (maximum input per embedding model)
- Pinecone: Chunking strategies for LLM applications
- LangChain: Text splitters (concept documentation)
- Unicode Standard Annex #29: Unicode Text Segmentation (sentence boundaries)
- MDN: Intl.Segmenter, locale-sensitive text segmentation in the browser
- Google AI for Developers: Long context in the Gemini API
- Anthropic: Context windows
- 3GPP TS 23.038: Alphabets and language-specific information (GSM 7-bit default alphabet)
- X Developer Platform: Counting characters
Frequently asked questions
How do I split text into chunks?
Choose the unit the destination counts (characters, words, tokens, sentences, paragraphs or lines), set the maximum per part, and cut on word or sentence boundaries rather than at the exact count. The Text Splitter does this in the browser and adds [1/N] labels, overlap, per-part copy and a .txt download.
How do I choose a chunk size for RAG?
Start from published defaults and test against your own questions. Microsoft’s chunking guidance gives a fixed size of, for example, 200 words or 600 characters with 10 to 15 percent overlap; OpenAI’s text-embedding-3 models accept at most 8,192 tokens per input. Specifications and tables want smaller chunks; narrative that needs cross-references wants larger ones.
What is the difference between a character text splitter and a recursive character text splitter?
A character splitter cuts at a fixed count of characters. A recursive character splitter, such as LangChain’s RecursiveCharacterTextSplitter, tries a list of separators in order, typically paragraph breaks, then line breaks, then spaces, so chunks follow the document’s structure as far as the size allows.
Why split by tokens instead of characters?
Because language models and embedding models measure input in tokens, and the character-to-token ratio changes with language, code and vocabulary. A part sized in tokens with the model’s own encoding fits by construction; one sized in characters is a guess.
What is chunk overlap and how much should I use?
Overlap repeats the end of one chunk at the start of the next so a fact on the boundary survives in at least one chunk. Microsoft’s guidance suggests roughly 10 to 15 percent of the content for retrieval; for text a person will read in order, use none.
Does splitting still matter with million-token context windows?
Yes. Google documents Gemini models with windows of 1 million or more tokens and Anthropic documents 200,000 and up to 1 million for Claude, but every token in the window is paid for, retrieval still returns chunks, and a model given one relevant paragraph usually answers better than one handed a hundred pages.
Keep reading
Written by SAVI. We build the tools we write about. Try the Text Splitter used in this post.