Study on Tokenization Differences in Chemical SMILES

Hunter Heidenreich · hf · 2026-07-08

This research discusses tokenization methods in chemical SMILES language models.

The authors found that Byte-pair encoding and Unigram-LM produce significantly different subword vocabularies, and they do not converge to the same result across different corpus types and vocabulary sizes.

Original post →

More from Research

Research channel →