What is a token in an LLM?
Large language models do not read characters or words. A tokenizer first splits text into tokens: pieces from a fixed vocabulary that the model learned to work with. Most current tokenizers use byte-pair encoding (BPE) or a close relative, with vocabularies of tens of thousands to a few hundred thousand entries.
Common English words are usually one token, often including the space before them. Rare words, names and typos are split into several pieces. Numbers are cut into short chunks, punctuation and indentation take tokens of their own, and text in scripts that were rarer in the tokenizer's training data is broken into many small pieces.
Tokens matter because everything is measured in them: the context window a model can see, the maximum length of its reply, rate limits and the price of each request.