Tokenization: The Hidden Layer Causing 54x AI Cost Gaps Across Languages
Google's 32-author survey on tokenization reveals it as the most underestimated component in LLMs, showing 94% of models reuse identical tokenizers, a 54.3x cost disparity between English and low-resource languages, and that BPE dominates generative models while fundamental fairness and efficiency challenges remain largely unsolved.
