The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
QUESTION — How can subword tokenization handle orthographic and structural variations without fragmenting the embedding space or losing information?
This paper presents The Functionalizer, a lossless pre-tokenizer framework that factors orthographic and structural variations into a compositional opcode/operand prefix stream prior to tokenization. Using parameterized transformation operators encoded in the Unicode Private Use Area for casing, diacritics, and character repetition, the framework achieves full corpus coverage with significantly smaller vocabularies. Downstream evaluations on 98M-parameter GPT-2 models demonstrate that it improves Python code syntax validity while reducing duplicate n-gram repetition in natural language prose.
It reduces actual vocabulary slot requirements by up to 19.7%.
It improves Python code syntax validity (9.12% vs. 7.70%) on 98M-parameter GPT-2 models.
mrkwanzaa · 18 Sept 2026
read the original ↗