The llama.cpp project released build b11455, introducing support for the PLaMo-3 tokenizer's pre-segmentation logic. This change implements hard boundaries before running the Unigram DP, specifically around specific text markers and runs of identical characters or spaces.
- Added vocab type "plamo3" to src/llama-vocab.cpp to handle the new tokenization rules.
- Reproduced the reference implementation's re.sub() passes by encoding each segment independently.
- Ensures code indentation and repeated punctuation are tokenized consistently with the PLaMo-3 reference.
This update allows llama.cpp to correctly tokenize inputs for PLaMo-3 models, matching the behavior of the original tokenizer.