A new preprint examines a basic but often overlooked step in large language model pipelines: tokenization. The authors note that different tokenizers can turn the same input text into token sequences of very different lengths, which can affect computational cost and, potentially, model behaviour. The paper frames tokenization as an optimisation problem, introducing counting-based and min-cost encoding strategies to control sequence length.
The proposed approach does not alter the model architecture; instead, it changes how text is mapped to tokens. By treating tokenization as a cost-minimisation task, the authors aim to reduce the number of tokens needed to represent a given string. The preprint does not claim a single best tokenizer, but rather suggests that the choice of encoding strategy matters and can be made more systematic.
Because this summary is based on a single preprint, there are no independent sources to compare for agreement or disagreement. The findings should be read as a proposal for further evaluation rather than a settled result.