Skip to content

deployment

SentencePiece

An unsupervised text tokenizer and detokenizer by Google that treats the input as a raw stream of characters, making it language-agnostic. SentencePiece can use BPE or unigram algorithms and is used by many multilingual models.

In practice

SentencePiece can tokenize Japanese, Chinese, and English text uniformly without requiring pre-segmentation.