Why LLMs Can't Count Letters in 'Strawberry': The Tokenization Problem Explained

This article explains how tokenization works in large language models, why it causes failures like miscounting letters in 'strawberry', details the BPE algorithm, discusses costs like language inequality and security risks, and explores alternatives like byte-level and dynamic tokenization.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Why LLMs Can't Count Letters in 'Strawberry': The Tokenization Problem Explained

Introduction: The Strawberry Problem

If you ask a large language model how many letters 'r' are in the word "strawberry", many state-of-the-art models answer incorrectly. The root cause is not model intelligence but how text is preprocessed before the model sees it. A 154-page survey led by Google with 32 authors from 20+ institutions places this failure front and center, arguing the problem lies in the tokenization step — the way text is split into numbered chunks called tokens.

1. Models Read Tokens, Not Characters

Humans see "strawberry" as ten distinct letters. Models, however, receive a sequence of integer IDs produced by a tokenizer. Common words like "strawberry" are often packaged as a single token, so the model never sees individual letters. When asked to replace each 'r' with 'z', the model outputs garbled results like "sztawbezzy" because the familiar token is broken into unfamiliar fragments.

Even place names differ: "California" becomes 10 bytes at the byte level, "Japan" 5 bytes. In early word-level models, missing vocabulary entries (e.g., "Americans" absent while "America" present) forced the tokenizer to emit a special <OOV> token, destroying semantic information.

Minor surface changes — adding quotes, a trailing space, or using a related form like "intelligent" — yield completely different token ID sequences. The model never reads raw text; it only processes the tokenizer's numeric output.

2. How Tokens Are Created: Byte-Pair Encoding (BPE)

Token vocabularies are not hand-crafted by linguists; they are learned statistically from massive corpora. The dominant algorithm is BPE (Byte-Pair Encoding). Starting from single characters/bytes, BPE repeatedly merges the most frequent adjacent pair into a new token until the vocabulary reaches a target size.

Worked example (from the article): Training data contains "low" (5×) and "lower" (2×). Initial tokens: l, o, w, e, r.

"l"+"o" co-occur 7× → merge into "lo".

"lo"+"w" co-occur 7× → merge into "low".

"low"+"e" co-occur 2× → merge into "lowe", etc.

This greedy merging is fast but not globally optimal; the survey notes (Section 13.2) that non-greedy merging can produce shorter sequences. Other algorithms include Unigram (pruning from a large candidate set) and WordPiece (merging to maximize next-token prediction likelihood, used in BERT).

A key insight: the same word can be segmented in exponentially many ways (e.g., "disjointed" → "dis"+"joint"+"ed" or a single token). The segmentation a model sees depends entirely on which tokenizer it was trained with.

3. The Costs of Token Packing

Capability: Arithmetic and Character-Level Tasks Suffer

GPT-2 tokenized numbers haphazardly (e.g., 8675309 → [8] [67] [5] [30] [9]). Newer models use a pre-tokenization rule grouping digits in threes from the left ( [867] [530] [9]), which contradicts right-to-left human arithmetic. Probe experiments show models only partially perceive sub-token structure. Counterfactual experiments (shuffling tokenization rules or using byte-level models) confirm that packing digits into tokens directly harms arithmetic and letter-counting performance.

Economic: Language Inequality

API pricing is per token. English, with abundant training data, gets compact tokens (few tokens per word). Low-resource languages are split into many more tokens. Attention cost grows quadratically with sequence length. The survey's Figure 6.2 shows generating the Universal Declaration of Human Rights with GPT-4o costs a Myanmar worker 54.3× more working seconds than a US worker, combining token inflation and income disparity. Proposals like PA-BPE remain largely academic.

Security: Untrained Tokens and Invisible Attacks

BPE's frequency-based vocabulary includes rare tokens (e.g., "SolidGoldMagikarp") that appear in tokenizer training data but almost never in model training data. Their embeddings stay near-random, creating hidden attack vectors. Adversaries can also swap visually identical characters (Latin 'a' vs. Cyrillic 'а') or insert zero-width joiners, causing the tokenizer to fragment safety-critical words and bypass content filters.

Ecosystem Concentration

A Hugging Face scan (Nov 2025) of ~740k models found 94% share byte-for-byte identical tokenizers with another model. The most downloaded tokenizer is BERT's (2018), accounting for 23.2% of 27.9 billion total downloads. Most teams never retrain tokenizers.

Anthropic's Claude bucks the trend: community reverse-engineering suggests its vocabulary shrank from ~49k (Claude 3–4.6) to ~16k (Claude 4.7+), possibly using dynamic programming for segmentation. The paper does not confirm a causal link to security.

4. Can We Change How Models Read?

Byte-level models (e.g., ByT5) use a 256-token vocabulary covering all languages. They excel at multilingual tasks but suffer from extremely long sequences (3× for Chinese), making inference impractically slow.

Dynamic tokenization (e.g., Meta's BLT) lets the model predict byte-level information density and cut at semantic boundaries, handing segmentation decisions to the model itself.

Visual tokenization (e.g., PIXEL, DeepSeek-OCR) renders text as images; a 22M-parameter PIXEL matches an 86M-parameter BERT on language understanding. The challenge is converting generated images back to editable text.

The survey authors conclude that while each new approach has merits, replacing the entrenched tokenization infrastructure demands massive retraining and evaluation investment. They highlight three priority directions: better tokenizer scaling laws, alternatives addressing subword defects, and more reliable evaluation benchmarks.

Conclusion

The "strawberry" error is not a memory lapse; it is a direct consequence of the tokenizer's packaging. Only 2% of top-tier ML conference abstracts mention tokenization, yet this preprocessing step silently determines every model's strengths and failure modes. Next time a model struggles with arithmetic, charges exorbitantly for non-English prompts, or crashes on gibberish, ask: what did the tokenizer actually feed into the model?

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

large language modelstokenizationNLPtokenizerBPEbyte-pair encodingLLM limitationstokenization survey
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.