How Large Language Models Work: From Prompt to Answer
A large language model splits your prompt into tokens, turns them into numbers, uses self-attention to work out context, then predicts the most likely next token. It repeats that prediction, one token at a time, until the answer is complete.

What a large language model is
A large language model (LLM) is an AI system that has learned patterns from a huge amount of text. Its core job is simple to state: given some text, predict the next token. Google’s machine learning course describes LLMs as, in effect, very powerful autocomplete systems that can predict thousands of tokens in a row.
What makes them “large” is the number of parameters, the adjustable numbers inside the network that are tuned during training. Google notes that today’s biggest models have hundreds of billions or even trillions of them.
Step 1: tokens and numbers
Computers work with numbers, not letters, so the first step is tokenisation. The model’s tokeniser splits your text into tokens, which can be whole words, pieces of words or single characters. For English, Google’s guide puts one token at roughly four characters, or about three-quarters of a word, so 400 tokens is around 300 words. A long word such as “unbelievable” might be split into a few pieces, but the exact split depends on the model.
Each token has an ID number in the model’s vocabulary. That ID is then converted into an embedding: a long list of numbers that the network learns during training and that captures something about the token’s meaning.
Step 2: the Transformer and self-attention
Most modern LLMs are built on the Transformer, an architecture introduced by Google researchers in the 2017 paper Attention Is All You Need. Its key idea is self-attention: for each token, the model asks how much every other token in the text should affect its meaning.
Take the sentence “The auto-rickshaw could not climb the hill because it was too steep.” Self-attention helps the model link “it” to “hill” rather than to “auto-rickshaw”. Change “steep” to “old”, and the link shifts to the auto-rickshaw. Because every token can look at every other token, the model can use context from anywhere in your prompt.
Step 3: predicting one token at a time
After the Transformer layers have processed your prompt, the model produces a probability for every token in its vocabulary. For “The capital of India is”, the token for “New” would get a very high probability and “bicycle” almost none.
The model picks a token, usually one of the most likely, adds it to the text, and runs the whole process again to choose the next one. This loop repeats until the reply is complete or a length limit is reached. That is why chatbot answers often appear on screen a few words at a time.
How LLMs are trained
Training happens in stages. In pre-training, the model reads an enormous amount of text and repeatedly guesses tokens that have been hidden from it, adjusting its parameters a little after each guess. This needs huge amounts of computing power.
A pre-trained model is good at continuing text but not at following instructions. So developers fine-tune it. In OpenAI’s 2022 InstructGPT research, people first wrote example answers, then ranked different model outputs, and those rankings were used to train the model further. This is called reinforcement learning from human feedback (RLHF), and it made answers more truthful and less toxic.
Limits to remember
Because an LLM predicts plausible text rather than checking a database of facts, it can hallucinate: give answers that sound confident but are wrong. Like other machine learning models, it can also reflect biases in its training data. A student in Chennai using a chatbot to revise for a Class 10 exam should treat it like a helpful study partner and check key facts against the NCERT textbook.
Frequently asked questions
What is a token?
A token is the basic unit of text an LLM works with. It can be a whole word, part of a word, a single character or a punctuation mark. In English, one token is roughly four characters, or about three-quarters of a word.
Does an LLM understand what it is saying?
It has learned rich patterns in language, but it works by predicting likely next tokens, not by checking facts. That is why it can produce fluent answers that are still wrong.
Why do LLMs sometimes hallucinate?
The model’s goal is to produce text that is likely given the prompt, not text that is verified. When it lacks reliable patterns for a question, it can still produce a confident-sounding answer that is incorrect.
Where did the Transformer come from?
It was introduced in the 2017 paper “Attention Is All You Need” by Ashish Vaswani and colleagues at Google. It replaced older step-by-step designs with attention, which made training faster and easier to run in parallel.
Sources & methodology
Every fact is checked against the sources below. We write original explanations and draw original graphics; no figures are copied from textbooks. Spotted an error? See our corrections policy.
- Introduction to Large Language Models (Google for Developers (Machine Learning Crash Course), accessed 30 Sept 2026)
- LLMs: What’s a large language model? (Google for Developers (Machine Learning Crash Course), accessed 30 Sept 2026)
- Attention Is All You Need (Vaswani et al., 2017) (arXiv, accessed 30 Sept 2026)
- Training language models to follow instructions with human feedback (Ouyang et al., 2022) (arXiv, accessed 30 Sept 2026)
Reviewed by Claude (pending owner check) · Updated
Embed this on your site
Paste this HTML into your page. Keep the credit line: it is required by the CC BY-NC 4.0 license.
Related infographics
TechGit Commands Cheat Sheet: The Essentials for Beginners
The Git commands beginners use most: set up a repository with init or clone, save work with add and commit, build features on branches with switch and merge, share work with fetch, pull and push, and undo safely with restore.
History & GeographyTimeline of the Internet: From ARPANET to Today
The internet grew from ARPANET, a US research network that sent its first message in 1969. TCP/IP in 1983, the World Wide Web from 1989 and public access, reaching India in 1995, led to about 6 billion users by 2025.
FinanceCompound Interest Explained: How Money Grows Over Time
Compound interest is interest earned on your interest. Each year’s interest is added to the balance, so the next year’s interest is worked out on a bigger sum. Over long periods this makes money grow far faster than simple interest.
FinanceHow Credit Scores Work in India: What Moves Your Score
A credit score is a three-digit number, such as Experian India’s 300–900 score, built from your borrowing record. Paying on time and using little of your credit limit matter most; the age and mix of your accounts matter less.
ScienceHow Photosynthesis Works: From Sunlight to Sugar
Photosynthesis is how plants use light energy to turn carbon dioxide and water into glucose, releasing oxygen as a by-product. It happens in chloroplasts in two linked stages: the light-dependent reactions and the Calvin cycle.
ScienceHow Vaccines Work: Training Your Immune System
Vaccines work by imitating an infection. They show your immune system a harmless antigen from a germ, so it makes matching antibodies and memory cells that can fight the real germ quickly if it ever arrives.
Get new InfoGraphHub resources by email
One short email when new free resources go live. No spam, unsubscribe anytime.