Skip to content
InfoGraphHub

How Large Language Models Work: From Prompt to Answer

A large language model splits your prompt into tokens, turns them into numbers, uses self-attention to work out context, then predicts the most likely next token. It repeats that prediction, one token at a time, until the answer is complete.

Flowchart of how a large language model answers: prompt typed, text split into tokens, tokens turned into numbers, self-attention adds context, next token predicted, loop until finished, answer shown.
InfoGraphHub · infographhub.com · CC BY-NC 4.0

What a large language model is

A large language model (LLM) is an AI system that has learned patterns from a huge amount of text. Its core job is simple to state: given some text, predict the next token. Google’s machine learning course describes LLMs as, in effect, very powerful autocomplete systems that can predict thousands of tokens in a row.

What makes them “large” is the number of parameters, the adjustable numbers inside the network that are tuned during training. Google notes that today’s biggest models have hundreds of billions or even trillions of them.

Step 1: tokens and numbers

Computers work with numbers, not letters, so the first step is tokenisation. The model’s tokeniser splits your text into tokens, which can be whole words, pieces of words or single characters. For English, Google’s guide puts one token at roughly four characters, or about three-quarters of a word, so 400 tokens is around 300 words. A long word such as “unbelievable” might be split into a few pieces, but the exact split depends on the model.

Each token has an ID number in the model’s vocabulary. That ID is then converted into an embedding: a long list of numbers that the network learns during training and that captures something about the token’s meaning.

Step 2: the Transformer and self-attention

Most modern LLMs are built on the Transformer, an architecture introduced by Google researchers in the 2017 paper Attention Is All You Need. Its key idea is self-attention: for each token, the model asks how much every other token in the text should affect its meaning.

Take the sentence “The auto-rickshaw could not climb the hill because it was too steep.” Self-attention helps the model link “it” to “hill” rather than to “auto-rickshaw”. Change “steep” to “old”, and the link shifts to the auto-rickshaw. Because every token can look at every other token, the model can use context from anywhere in your prompt.

Step 3: predicting one token at a time

After the Transformer layers have processed your prompt, the model produces a probability for every token in its vocabulary. For “The capital of India is”, the token for “New” would get a very high probability and “bicycle” almost none.

The model picks a token, usually one of the most likely, adds it to the text, and runs the whole process again to choose the next one. This loop repeats until the reply is complete or a length limit is reached. That is why chatbot answers often appear on screen a few words at a time.

How LLMs are trained

Training happens in stages. In pre-training, the model reads an enormous amount of text and repeatedly guesses tokens that have been hidden from it, adjusting its parameters a little after each guess. This needs huge amounts of computing power.

A pre-trained model is good at continuing text but not at following instructions. So developers fine-tune it. In OpenAI’s 2022 InstructGPT research, people first wrote example answers, then ranked different model outputs, and those rankings were used to train the model further. This is called reinforcement learning from human feedback (RLHF), and it made answers more truthful and less toxic.

Limits to remember

Because an LLM predicts plausible text rather than checking a database of facts, it can hallucinate: give answers that sound confident but are wrong. Like other machine learning models, it can also reflect biases in its training data. A student in Chennai using a chatbot to revise for a Class 10 exam should treat it like a helpful study partner and check key facts against the NCERT textbook.

Frequently asked questions

What is a token?

A token is the basic unit of text an LLM works with. It can be a whole word, part of a word, a single character or a punctuation mark. In English, one token is roughly four characters, or about three-quarters of a word.

Does an LLM understand what it is saying?

It has learned rich patterns in language, but it works by predicting likely next tokens, not by checking facts. That is why it can produce fluent answers that are still wrong.

Why do LLMs sometimes hallucinate?

The model’s goal is to produce text that is likely given the prompt, not text that is verified. When it lacks reliable patterns for a question, it can still produce a confident-sounding answer that is incorrect.

Where did the Transformer come from?

It was introduced in the 2017 paper “Attention Is All You Need” by Ashish Vaswani and colleagues at Google. It replaced older step-by-step designs with attention, which made training faster and easier to run in parallel.

Sources & methodology

Every fact is checked against the sources below. We write original explanations and draw original graphics; no figures are copied from textbooks. Spotted an error? See our corrections policy.

  1. Introduction to Large Language Models (Google for Developers (Machine Learning Crash Course), accessed 30 Sept 2026)
  2. LLMs: What’s a large language model? (Google for Developers (Machine Learning Crash Course), accessed 30 Sept 2026)
  3. Attention Is All You Need (Vaswani et al., 2017) (arXiv, accessed 30 Sept 2026)
  4. Training language models to follow instructions with human feedback (Ouyang et al., 2022) (arXiv, accessed 30 Sept 2026)

Reviewed by Claude (pending owner check) · Updated

Embed this on your site

Paste this HTML into your page. Keep the credit line: it is required by the CC BY-NC 4.0 license.

Get new InfoGraphHub resources by email

One short email when new free resources go live. No spam, unsubscribe anytime.