How Large Language Models Work
Understanding the Technology You Will Build With
You do not need to be an AI researcher to build AI automation systems. But having a solid conceptual understanding of how LLMs work makes you a much better builder. It helps you:
- Write better prompts
- Understand why models succeed and fail
- Set realistic expectations with stakeholders
- Debug AI-powered systems effectively
What is a Language Model?
A language model is a system trained to predict the probability of sequences of words (or tokens). Given a sequence of text, it estimates what comes next.
Early language models were simple: they looked at the last two or three words and predicted the next one based on statistics from a training corpus. Modern Large Language Models are vastly more sophisticated, but the fundamental idea remains: predict what comes next.
Tokens: The Unit of Language
LLMs do not process text character by character or word by word. They process tokens -- chunks of text that are typically three to four characters long.
"Hello, how are you today?"
Tokens: ["Hello", ",", " how", " are", " you", " today", "?"]
A rough rule of thumb: 1 token is approximately 0.75 words, or about 4 characters. The text "I love building AI systems" is approximately 6 tokens.
Why does this matter? Because LLMs have a context window -- a maximum number of tokens they can process at once. GPT-4's context window is 128,000 tokens (about 100,000 words). Claude's is 200,000. This determines how much text an LLM can read and reason about in one interaction.
How LLMs Are Trained
Training a large language model involves three stages:
Stage 1: Pre-training
The model is trained on an enormous dataset of text -- trillions of tokens scraped from the internet, books, code repositories, and other sources. The training objective is simple: predict the next token.
Through this process, the model implicitly learns:
- Grammar and syntax
- Facts about the world
- Reasoning patterns
- Code structure and semantics
- Relationships between concepts
This is why LLMs can answer questions, write code, and translate languages without being explicitly programmed to do so.
Stage 2: Supervised Fine-tuning (SFT)
The pre-trained model knows how to predict text, but it does not know how to be helpful. In this stage, human contractors create examples of good conversations (question and ideal answer pairs) and the model is fine-tuned on these examples.
Stage 3: Reinforcement Learning from Human Feedback (RLHF)
Human raters compare different model outputs and indicate which are better. This feedback trains a reward model that the LLM is then optimised against. This stage is what makes models like GPT-4 and Claude genuinely helpful, harmless, and honest rather than just statistically plausible.
The Transformer Architecture
The architectural breakthrough behind modern LLMs is the Transformer, introduced by Google researchers in 2017 in a paper called "Attention Is All You Need."
The key innovation was the attention mechanism: the ability for every token to attend to every other token in the context and weight their relevance. This lets the model understand long-range dependencies in text.
For example, in the sentence "The bank by the river was steep, so I climbed it", the word "it" refers to "bank," not "river." Attention lets the model resolve this ambiguity by considering the full sentence.
You do not need to implement this. But understanding that transformers reason about the relationships between all tokens in the context helps you understand why:
- Longer, more detailed prompts often produce better results
- Context at the beginning of the prompt matters
- Models can lose track of information in very long contexts
What LLMs Are Good At
| Capability | Examples |
|---|---|
| Text generation | Writing emails, blog posts, code, creative content |
| Summarisation | Condensing long documents into key points |
| Classification | Sorting text into categories (sentiment, topic, intent) |
| Extraction | Pulling structured data from unstructured text |
| Translation | Converting text between languages |
| Question answering | Answering questions about provided documents |
| Reasoning | Breaking down problems step by step |
| Code generation | Writing, explaining, and debugging code |
What LLMs Are Not Good At
Understanding limitations is as important as understanding capabilities:
Hallucination: LLMs can generate confident-sounding but factually incorrect information. They do not "know" facts the way a database does -- they pattern-match to likely text. Always verify critical facts.
Arithmetic: LLMs are poor at precise calculation. Use code for maths.
Recency: LLMs have a training cutoff date and do not know about events after that date. Use retrieval systems to provide current information.
Consistent identity: LLMs do not have memory between separate conversations unless you explicitly provide previous context.
Determinism: The same input can produce different outputs. LLMs are probabilistic, not deterministic. (You can reduce this with the temperature parameter.)
Key Parameters You Will Use
When calling an LLM API, you can control several parameters:
Temperature (0 to 2): Controls randomness. 0 = deterministic, always the most likely next token. 1 = normal randomness. 2 = very creative and unpredictable. For production automations requiring consistency, use temperature 0 or close to it.
Max tokens: The maximum number of tokens in the response. Limits cost and response length.
System prompt: Instructions given to the model before the user's input. This is where you define the model's role, constraints, and output format.
Top-p (0 to 1): An alternative to temperature for controlling randomness. Usually only adjust one of temperature or top-p.
Try it yourself
Key Takeaways
- LLMs predict the next token in a sequence; through massive pre-training, they implicitly learn facts, reasoning, and language.
- The context window defines how much text an LLM can process at once, ranging from tens of thousands to hundreds of thousands of tokens.
- Training has three stages: pre-training on vast text data, supervised fine-tuning on good examples, and RLHF to align with human preferences.
- LLMs excel at text generation, summarisation, classification, extraction, and code generation but struggle with arithmetic, recency, and consistency.
- Temperature controls randomness in output. Use low temperature (near 0) for consistent production automations and higher values for creative tasks.
Quick Quiz
1.What is a token in the context of large language models?
2.What is 'hallucination' in the context of large language models?
3.What effect does setting temperature to 0 have on an LLM's output?
4.What was the key innovation of the Transformer architecture?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx