A machine cannot read words. Text is chopped into pieces, and each piece becomes a number. Everything downstream — cost, memory, even accuracy — is decided by how well that chopping goes.
A box of LEGO bricks
Imagine you can only speak using a fixed box of LEGO bricks. A common word like weather has its own brick: you just grab it. An unusual word has no brick of its own, so you build it out of several smaller ones.
That box is the model’s vocabulary, typically 32,000 to 200,000 bricks. Every piece of text you send gets built out of those bricks, and each brick has a number.
Here’s what actually happened when I fed two words into a well-known model:
| Word | Pieces | |
|---|---|---|
weather |
weather |
one piece |
hyperlipidemia |
hyper · lip · id · emia |
four pieces |
Notice something odd. A person would split that word as hyper–lipid–emia. The machine split it as hyper–lip–id–emia. That lip is the lip on your face — nothing to do with fats in the blood.
The machine didn’t choose pieces that carry meaning. It chose pieces that are common on the internet. And the internet talks about lips far more than lipids.
Count what it costs
Specialist language shatters into more pieces than everyday language. I measured pieces per word on the same model, for two kinds of text:
Nobody designed this vocabulary
The vocabulary is learned from a corpus, usually with byte-pair encoding (BPE): start with individual characters, then repeatedly merge the most frequent adjacent pair into a new token. Do that 50,000 times and you have a vocabulary.
The consequence is the whole chapter in one line: frequency in the training corpus, not meaning, decides what gets its own token.
What specialist language costs
Same two models, two kinds of text:
| Everyday English | Clinical notes | Penalty for medicine | |
|---|---|---|---|
| General model | 1.08 | 1.97 | +82% |
| Medical model | 1.17 | 1.37 | +17% |
A medical-domain tokenizer cuts the cost of clinical text by about 30%, and pays for it by being slightly worse at everyday English. That’s not a flaw — it’s a choice about which world the model lives in.
thrombocytopenia
general model → th · rom · b · ocy · top · en · ia (7 pieces)
medical model → thrombocytopenia (1 piece)
Seven fragments, or one whole idea. That’s a dangerously low platelet count — exactly the kind of thing you want a medical system to understand as one concept.
I looked inside both vocabularies. The general model reserved a slot for harbaugh, an American football coach’s surname. The medical model spent its slot on leishmaniasis, a tropical parasitic disease. Neither is wrong. But you wouldn’t want the football dictionary reading your blood report.
The three bills
Money. These systems bill per token. A 400-word clinical note is ~790 tokens instead of ~430. Every note, forever.
Memory. A model holds a fixed number of tokens at once. Whatever that limit is, you fit roughly half as much patient history when the text is specialist.
Meaning. This is the one that hurts. The model isn’t reading “hyperlipidemia” — it’s reading four unrelated fragments and rebuilding the idea from scratch, every single time it appears.
Where it breaks
- You can’t swap in a better tokenizer. The model’s entire learned memory is indexed by those token numbers. Change the numbering and every word points at the wrong meaning. A specialist vocabulary means retraining — a real project, not a setting.
- Numbers and code tokenize badly.
2026may be one token or three; that’s part of why models fumble arithmetic. - Non-English text pays a tax. Languages under-represented in the training corpus need far more tokens per word — sometimes 2–3×, for identical meaning. The same API call costs more in Hindi than in English.
- Token boundaries hide characters. “How many r’s in strawberry?” is hard partly because the model sees two or three chunks, not eleven letters.
Whiteboard check
Marker in hand, someone watching. Could you get through these without notes?
NoteWhat is BPE and why is it used instead of word-level or character-level tokenization?
Byte-pair encoding starts from characters and repeatedly merges the most frequent adjacent pair. Word-level vocabularies can’t handle unseen words and get enormous; character-level sequences are far too long for attention, which costs quadratic time. BPE sits in between: a fixed-size vocabulary, no out-of-vocabulary words (anything can be spelled out from smaller pieces), and sequences short enough to process.
NoteYour RAG system over clinical documents is slower and pricier than expected. Where do you look first?
Tokens per word. Specialist text can nearly double the token count of the same content, which inflates embedding cost, retrieval chunk sizes and prompt length all at once. Measure the ratio on your own corpus before assuming the model or the retriever is at fault.
NoteWould a domain-specific tokenizer fix a model that’s bad at medical text?
It fixes the representation problem, not the knowledge problem, and only if you retrain or continue pretraining — the vocabulary can’t be swapped into an existing model. Expect a ~30% reduction in tokens for clinical text, and slightly worse performance on general text.
TL;DR
- Text becomes pieces, and pieces become numbers. Pieces are chosen by frequency, not meaning.
- Specialist language shatters worst: +82% tokens for clinical text on a general model.
- A domain tokenizer cuts that to +17%, but it must be baked in from the start.
- Cost, context limits and accuracy all trace back to this one step.