Glossary

Tokenization

The process of splitting text into tokens — the basic units an LLM processes — which are roughly word pieces, not whole words, explaining why a 10,000-word document uses more than 10,000 tokens.


What it means

Tokenization is the process of converting text into the discrete units — called tokens — that a large language model actually processes. Tokens are not whole words: they are word pieces, chosen by an algorithm that balances vocabulary size against coverage of rare words and scientific terminology.

Rules of thumb for estimating tokens:

  • Common English words: typically 1 token each
  • Longer or less-common words: 2–3 tokens (e.g., “biosynthesis” → “bio” + “synthesis” = 2 tokens)
  • Chemical names, SMILES strings, gene identifiers: can be many tokens per term
  • A typical academic paper (~8,000 words): roughly 11,000–13,000 tokens
  • The rough average: 1 token ≈ 0.75 words, or 100 tokens ≈ 75 words

Why it matters for researchers

Scientific text uses more tokens than general text. A methods section dense with chemical formulas, gene names, and species nomenclature may tokenize at closer to 1:1 word-to-token ratio than the 0.75 average. This means your document may hit a model’s context limit faster than expected if it contains a lot of specialized terminology.

Token limits affect pricing too. LLM API pricing is per token (both input and output). If you’re building a research pipeline that processes many documents, estimating token counts accurately matters for cost forecasting.

Why some terms are expensive to tokenize: Models learn their tokenizer vocabularies from training corpora. Words that appear rarely in the training data get broken into more subword pieces. Highly specialized scientific terminology — IUPAC chemical names, uncommon gene symbols, journal-specific abbreviations — often tokenizes inefficiently.

Practical tools: The Tiktokenizer website (by OpenAI) and similar tools let you paste text and see exactly how a specific model’s tokenizer splits it before sending it to the API.