2. SMILES¶
pip install SmilesPE
from SmilesPE.pretokenizer import atomwise_tokenizer
smi = 'CC[N+](C)(C)Cc1ccccc1Br'
toks = atomwise_tokenizer(smi)
print(toks)
K-mer Tokenzier
from SmilesPE.pretokenizer import kmer_tokenizer
smi = 'CC[N+](C)(C)Cc1ccccc1Br'
toks = kmer_tokenizer(smi, ngram=4)
print(toks)