How An LLM Reads And Writes: Tokens, Embeddings And Next-Token Prediction
Large language models turn text into numbered pieces, turn those into vectors and then predict one piece at a time. Here is that loop explained step by step.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
A large language model (LLM) can draft a contract clause, summarise a report or answer a question in fluent prose. Underneath, it does one narrow thing over and over: it looks at the text so far and estimates which small piece of text is likely to come next.
Understanding that loop explains a lot about how these systems behave, including why they are good with language, why their costs are measured in tokens, and why they sometimes state false things with complete confidence. This article covers the three ideas you need: tokens, embeddings and next-token prediction.
Tokens: How Text Becomes Pieces
A model cannot read letters or words directly. Its first step is to split text into tokens, which are pieces drawn from a fixed vocabulary. That vocabulary is built by a separate tool, the tokeniser, which is trained on sample text before the model itself is trained. Common words are often a single token. Rarer words are broken into smaller pieces.
For example, a tokeniser might split “unreconciled invoices” into pieces such as “un”, “reconc”, “iled” and “ invoices”. The exact split depends on the model, so treat this as an illustration. Each piece then becomes a number, its position in the vocabulary.
Sennrich, Haddow and Birch made the case for this approach in a 2016 paper on machine translation. Their insight was that a translation system can handle a word it has never met as a whole if it can build that word from smaller fragments it does know, and one of the ways they chose those fragments was a method called byte pair encoding.1 The practical benefit is that a vocabulary of tens of thousands of pieces can cover almost any word, name or product code, although a character the tokeniser never saw in training can still cause problems.
Embeddings: How Pieces Become Meaning
A token number on its own carries no meaning; token 4,512 is not “bigger” than token 12. So the model converts each token into an embedding: a long list of numbers, often hundreds or thousands of them. Each number counts as one dimension, so the list places the token at a point in a space with that many dimensions.
These positions are learned during training. Tokens used in similar contexts tend to end up close together. As an illustration, you would expect the vectors for “invoice” and “bill” to sit nearer to each other than either does to “giraffe”. Mikolov and colleagues showed in 2013 that useful word vectors of this kind could be learned efficiently from very large amounts of text.4 Modern LLMs learn their embeddings as part of the whole model, and the layers above refine each vector using the surrounding words, so “bank” in “river bank” ends up represented differently from “bank” in “bank transfer”. Researchers showed in 2018 that representations shaped by context in this way capture the different senses a word can have.5
Embeddings are useful beyond the model itself. Search systems can convert documents and questions into embeddings and look for the nearest matches, which is the basis of retrieval-augmented generation, covered later in this group.6
Next-Token Prediction: The Core Loop
Once the input is a sequence of vectors, the model’s layers process it and produce a score for every token in its vocabulary. Those scores are turned into probabilities that add up to one, as in the original transformer design.7 The model then picks a token, adds it to the text, and repeats the whole process with the slightly longer text.
- Split Into Tokens
The prompt "Please pay the invoice by" becomes a short list of token numbers.
- Look Up Embeddings
Each token number is swapped for its learned vector.
- Process Through The Layers
The model combines information across all the tokens so far.
- Score Every Possible Next Token
The output is a probability for each token in the vocabulary.
- Choose One Token
For example " Friday", picked by the decoding settings.
- Append And Repeat
The chosen token joins the input and the loop runs again until the reply ends.
The model learned these probabilities during pretraining, where it read enormous amounts of text and was repeatedly asked to predict the next token, adjusting its parameters each time it was wrong. Parameters are the adjustable numbers inside the model, and a large model has billions of them. GPT-3, described in 2020 with 175 billion parameters, is a well-documented example of this approach and showed that a large enough next-token predictor could perform many tasks from just a few examples in the prompt.8 How the model is then shaped into a helpful assistant is the subject of From Raw Model To Assistant.
What This Explains About LLM Behaviour
Seeing the model as a next-token predictor clears up several common puzzles.
It writes fluent text because fluent text is exactly what it was trained to continue. It can be wrong while sounding certain, because the loop picks plausible continuations; nothing in the basic loop checks them against a source of truth. It can give different answers to the same question because the choice at each step can involve chance, controlled by settings such as temperature.9 It also often struggles with tasks like counting the letters in a word, and a likely reason is that it works with tokens rather than individual letters.
None of this means the output is random or useless. Next-token prediction at scale captures a great deal of grammar, facts and reasoning patterns found in text. It does mean that an LLM’s output should be treated as a well-informed draft rather than a verified record. The final article in this group, Using LLMs Well, covers how to reduce errors in practice.
Footnotes
-
R. Sennrich, B. Haddow and A. Birch, “Neural Machine Translation of Rare Words with Subword Units”, ACL 2016, arXiv:1508.07909. arxiv.org ↩
-
Anthropic (vendor documentation), “Pricing”, accessed 7 October 2026. platform.claude.com ↩
-
A. Petrov, E. La Malfa, P. H. S. Torr and A. Bibi, “Language Model Tokenizers Introduce Unfairness Between Languages”, NeurIPS 2023, arXiv:2305.15425. arxiv.org ↩
-
T. Mikolov, K. Chen, G. Corrado and J. Dean, “Efficient Estimation of Word Representations in Vector Space”, arXiv:1301.3781, January 2013. arxiv.org ↩
-
M. E. Peters et al., “Deep contextualized word representations”, NAACL 2018, arXiv:1802.05365. arxiv.org ↩
-
P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, NeurIPS 2020, arXiv:2005.11401. arxiv.org ↩
-
A. Vaswani et al., “Attention Is All You Need”, arXiv:1706.03762, June 2017, section 3.4. arxiv.org ↩
-
T. B. Brown et al., “Language Models are Few-Shot Learners”, arXiv:2005.14165, May 2020. arxiv.org ↩
-
A. Holtzman, J. Buys, L. Du, M. Forbes and Y. Choi, “The Curious Case of Neural Text Degeneration”, ICLR 2020, arXiv:1904.09751. arxiv.org ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.