Researchers propose Counting and Filtering (CNF) for vocabulary construction and Min-Cost Encoding (MCE) for text segmentation to enhance tokenization in large language models. This combination offers higher token efficiency, greater scalability, and lower dependency compared to Byte Pair Encoding (BPE).

  • CNF builds a raw vocabulary by counting valid substrings and filters it based on actual usage during MCE segmentation.
  • MCE minimizes a cost function over text segments to determine the optimal global segmentation without relying on merge lists or token probabilities.
  • With a 250K vocabulary, CNF-MCE increases compression rate by 26% and 30% on English web text compared to o200k_base and qwen250k tokenizers.
  • Scaling the vocabulary to 1M entries yields over 60% token efficiency improvement and raises vocabulary utilization from 52.9% to 96.9%.
  • Language models trained at 1.8B and 8B scales achieve comparable performance to BPE-based models across 11 benchmarks.

The results demonstrate that CNF-MCE significantly improves token efficiency while maintaining competitive downstream performance, applicable to vocabularies built from various methods.