Tokenization: The Hidden Layer Causing 54x AI Cost Gaps Across Languages
Google's 32-author survey on tokenization reveals it as the most underestimated component in LLMs, showing 94% of models reuse identical tokenizers, a 54.3x cost disparity between English and low-resource languages, and that BPE dominates generative models while fundamental fairness and efficiency challenges remain largely unsolved.
Tokenization: The Overlooked Foundation of Modern LLMs
The article opens with the now-famous "strawberry" test: even advanced models miscount the letter "r" in "strawberry" because they never see characters — they see tokens. The token "strawberry" may be split as str + aw + berry, so the model has no access to individual letters. This illustrates the core thesis: tokenization, the very first step in the LLM pipeline, is the most underestimated component, receiving less than 2% of top-conference paper attention (NeurIPS, ICLR, ICML, CL) since 2020 despite being the sole interface between raw text and model internals.
How a Tokenizer Is Built: A Five-Stage Pipeline
The survey breaks tokenizer construction into a clear pipeline:
Corpus + initial alphabet → Normalization (e.g., Unicode NFC, lowercasing) → Pre-tokenization (splitting on whitespace/punctuation) → Training algorithm (the core) → Usable tokenizer .
At inference, text passes through normalization, pre-tokenization, and encoding into a token sequence; model outputs are decoded via the inverse process (detokenization).
The Three Dominant Algorithm Families
The survey provides a complete genealogy of modern tokenization algorithms, highlighting three main families:
BPE (Byte Pair Encoding) — Bottom-up. Starts from single bytes, repeatedly merges the most frequent adjacent token pair. Originated as a 1994 compression algorithm, adapted for NLP in 2016, now the absolute mainstream for generative LLMs.
Unigram — Top-down. Begins with a huge vocabulary, then prunes tokens by likelihood, discarding those whose removal hurts probability least.
WordPiece — BPE's close relative; merges based on language-model score rather than raw frequency. Uses MaxMatch greedy longest-match decoding at inference. Still used by the BERT family for classification tasks.
Numerous variants exist (BPE-Dropout adds stochasticity, SaGe introduces contextual scoring, PathPiece builds a graph of all possible segmentations to find the shortest path), but the survey notes these remain rarely adopted at scale — not because they are inferior, but because replacing the tokenizer foundation is prohibitively expensive.
Language Fairness: The 54.3× Cost Gap
Vocabulary allocation follows corpus frequency: English gets short, efficient tokens; low-resource languages are fragmented into many tokens. Since APIs charge per token, this creates massive cost inequality. The survey calculates the cost of generating the Universal Declaration of Human Rights with GPT-4o at $10 per million tokens, converted into average hourly wages:
US vs. Myanmar users: 54.3× difference.
This is framed as a fairness issue, not just a technical detail. Research proposals like PA-BPE (prioritizing the worst-compressed language during training) exist but remain largely on paper.
Hugging Face Census: 94% of Models Share Identical Tokenizers
The survey team scanned 739,147 models on Hugging Face, successfully loading 57.3% and identifying 24,798 distinct tokenizers . Key findings:
94% of models use a tokenizer that is byte-for-byte identical to another model's tokenizer.
Top 10 tokenizers cover 32% of all models; Llama 3.1 alone accounts for 6.9%.
The most downloaded tokenizer is the 2018 BERT tokenizer, representing 23.2% of 279 billion total downloads.
For frontier closed-source models, vocab sizes cluster at 100–275k tokens , almost exclusively BPE. The notable exception is Claude : its tokenizer details are unpublished, but community reverse-engineering suggests it uses a shortest-path algorithm (not standard BPE) and its vocabulary shrank from ~49k to ~16k in the latest generation — the only major vendor deviating from the BPE norm.
Trend analysis reveals a clear division of labor post-2023: BPE dominates generative models, WordPiece remains entrenched in BERT-style classification . Average vocabulary size surged after 2023, and non-Latin token share rebounded — driven by multilingual vocabularies from Chinese vendors such as Qwen and DeepSeek.
Future Directions: Beyond Static Vocabularies
The survey concludes with an explicit "undeveloped map" of promising directions:
Byte/character-level models (e.g., hierarchical byte-level LMs) to eliminate the vocabulary bottleneck entirely.
Dynamic tokenization that adapts segmentation to context.
Visual tokenization for multimodal inputs.
The 32 authors state plainly: "There is still vast room for improvement." A survey does not solve problems, but it maps every crack in the foundation — when 2% of research attention meets 100% of system dependency, what is underestimated is not a single algorithm but the entire stage.
Full paper:
https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlpSigned-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
