
Well, there’s no hiding from it. LLMs are here to take our jobs, and possibly our wives and pets. In an effort to know thy enemy before it destroys us, let’s try to figure out how this stuff works.
Alright, time’s a-wasting. Let’s get started
Buzzword Breakdown
There are a lot of words out there. So let’s define some key terms before we go any further. These go roughly from most broad to most specific.
- Artificial Intelligence (AI): The big umbrella term. AI refers to any system that attempts to mimic human intelligence to perform tasks, whether that’s recognizing speech, playing chess, or generating memes. Not all AI is smart; some of it is just smoke and mirrors. But if it tries to “act human,” it probably falls under this category.
- Machine Learning (ML): A subfield within the field of AI. Machine learning is what happens when we stop hard-coding rules and instead let algorithms learn patterns and rules directly from data. It’s less “tell the computer what to do” and more “teach the computer to figure it out.”
- Neural Network: A neural network is a system of connected “neurons” (basically just mathematical functions interacting with each other). You feed data into the first layer of the neural network, a layer being a group of neurons that work together to transform input into output, it passes through a bunch of other layers, and you get some kind of output at the end. With enough layers and data, a neural network can produce outputs that recognize incredibly subtle patterns.
- Deep Learning: A subset of machine learning. Deep learning uses neural networks to model complex patterns in data. It’s especially good at tasks like image recognition, voice synthesis, and language understanding. It’s called “deep” because the neural networks required for deep learning often have many, many layers.
- Natural Language Processing (NLP): A field focused on enabling machines to understand and work with human language. NLP covers everything from translating French to English, to analyzing the tone of a tweet, to auto-completing your email.
- Large Language Model (LLM): A specific kind of deep learning model designed for NLP. An LLM is trained on massive amounts of text, books, articles, tweets, Reddit comments, whatever, and learns statistical patterns in language. Once trained, it can generate new text, answer questions, write code, and much more. LLMs are basically what you get when you apply deep learning to huge piles of language data.
What Even Is an LLM, Anyway?
First things first: LLM stands for “Large Language Model.” At its core, it’s just a machine learning model trained on a massive amount of data. These models learn patterns in that data to do things like read, understand, and respond to text, images, audio, or other inputs.
Now, you might have heard of different kinds of LLMs: GPT, BERT, and so on. They’re all part of the same family, but they work a bit differently. For example, GPT is a type of LLM built mainly for generating text by predicting the next word in a sequence, one after another. BERT, on the other hand, is designed more for understanding text, which helps with things like search and classification. We will get into more details on these LLMs a bit further on.
Despite those differences, they’re all basically just pattern-recognition machines trained on huge piles of data. They don’t know whether something is right or wrong, and they definitely don’t possess “consciousness” unless you consider a giant matrix of numbers to be conscious. If you’re feeling philosophical, maybe that isn’t all that different from us humans, if you really think about it.
Further reading suggestion: You Look Like a Thing and I Love You by Janelle Shane.
A Brief History of LLMs
LLMs certainly feel like they came out of nowhere. One minute I’m sifting through StackOverflow threads, and the next I’m pasting a stack trace into an AI-powered IDE and saying “plz fix” like I’m talking to an intern polishing a PowerPoint. But these models didn’t magically appear, they’re the result of decades of research, and hardware that finally caught up to the theory.
AI research, as we know it, began in the 1950s with early systems that tried to mimic human reasoning using logic and rules, think theorem solvers and chess programs, like the ones discussed by these guys. These early projects were ambitious but brittle, and when they failed to deliver widespread success, funding dried up. This period of stalled progress in the world of AI became known as the AI Winter.
There were intermittent AI successes here and there, important milestones that pushed the field forward, but they weren’t LLMs. No deep learning or anything like that. Like in 2011, when IBM Watson won Jeopardy! against Ken Jennings and Brad Rutter, using a combination of keyword matching, database lookups, and early NLP techniques.
The real game-changer came in 2017, when Google researchers published the paper, Attention Is All You Need, introducing the transformer architecture. This architecture allowed models to process all inputs to a model at once using a mechanism called “self-attention”. It was fast, parallelizable, and shockingly good at understanding language structure. This solved key limitations with earlier models based on recurrent neural networks (RNNs), which processed text sequentially and often struggled to capture long-range dependencies or parallelize efficiently.
The transformer architecture inspired a flood of transformer-based LLMs. In 2015, OpenAI was founded with the mission of building safe, widely beneficial artificial general intelligence, releasing several models, including the popular ChatGPT, which is just an interface to the company’s underlying LLMs (like GPT-3.5 and GPT-4). Since then, everyone’s jumped in. Meta released LLaMA, Anthropic built Claude, Google launched Gemini, and Hugging Face turned models into downloadable APIs. LLMs stopped being a research novelty and became everyday tools, from coding assistants to customer support bots and everything in between.
LLMs didn’t come out of nowhere. But now that they’re here, they’re everywhere.
Tokenization
LLMs fundamentally process data as large sets of numbers. But LLMs doing NLP, almost by definition, deal with words, not numbers. So what’s going on here?
This brings us to one of the first critical concepts for understanding how an LLM works: tokenization.
Tokenization is the process of breaking text into smaller units called “tokens” and assigning a token ID (just a number) to each one. A simple example would be assigning a token to each word in the dictionary. In this way, any piece of text can be represented as a sequence of tokens.
When you hear about “token context” or “token response size” in LLMs, it is simply referring to the number of tokens the LLM will consider in your input or include in its response. So, if you request a 512-token response and each token corresponds to one word, you might get roughly a 512-word reply. In practice, tokens vary in size, as we’ll explore in future sections, but that’s the basic idea.
Byte-Pair Encoding and other Tokenization Methods
In practice, tokens are usually words, character sets, or combinations of words and punctuation. It’s rarely as simple as one character per token or even one word per token. Because what happens when you get a word you don’t recognize? Like some Dr. Seuss “Gluppity-glup” type nonsense? What do you do then, smart guy?
That’s where more advanced tokenization algorithms come in. One of the most popular is Byte-Pair Encoding, or BPE. BPE is an old-school algorithm from 1994 that’s found new life in LLM development. It’s a relatively simple algorithm that works by repeatedly replacing the most common contiguous sequences of characters with a placeholder token that points to that sequence in a lookup table.
Let’s say you have the word “walking.” A tokenizer might break it down like this:
- Instead of treating “walking” as one whole token, it could split it into smaller chunks like “walk” and “ing.”
- So, the tokenizer assigns a token ID to “walk” and another token ID to “ing”
- This way, if the model already knows “walk” and “ing” from other words, it can understand and generate new words it hasn’t seen before by combining tokens.
By operating on subwords rather than full words or phrases, BPE can efficiently generate tokens from any input without requiring prior knowledge of every word. This helps keep the total number of tokens manageable.
OpenAI maintains an open-source implementation of BPE called tiktoken, which is widely used by LLM developers. Other tokenization methods include WordPiece (used by BERT) and SentencePiece (used by LLaMA). These are similar to BPE in that they aim to create an efficient set of tokens for a text dataset, but their exact techniques differ slightly.
And if you’re thinking BERT is from Sesame Street and LLaMA sounds like one of those fuzzy goat-type animals, that’s totally fine - stick with us for a bit. It’ll all start to make sense.
For reference, OpenAI’s GPT-2 has a vocabulary size of 50, 257 tokens and a context window of 2,048 tokens. This means it recognizes about 50,000 unique token units in its training data and can process up to 2,048 tokens at once in prompts and responses. (Source: OpenAI GPT-2 paper).
Fun fact: “Gluppity-glup” is a real term from Dr. Seuss’s The Lorax.
Corpus
A corpus is a large collection of text or other data used to train a LLM. It can include various types of unstructured data such as text documents, speech transcripts, videos, or anything else depending on the model’s intended capabilities. For multimodal models, the corpus may contain some or all of these data types.
Collecting high-quality training data, and a lot of it, is one of the most important steps when preparing to train an LLM. Common sources include datasets like Common Crawl, public code repositories, or maybe even all the text messages you’ve ever sent that might be sitting on some NSA server somewhere, if you’re into that kind of thing.
Cleaning this data is crucial for building a good model. As the saying goes, garbage in, garbage out.
Word Embeddings
Word embeddings are vector representations of words in a high-dimensional space, where semantically similar words are positioned closer together.
Now, that sentence is a doozy. I had to just get it out there, but don’t worry, we’re going to walk it back a bit. That’s my main gripe with a lot of books and articles on LLMs. They just drop these dense sentences on you and then cruise right past while your eyes glaze over and you start wondering if you should’ve just become a real estate agent or something.
Anyway, back to word embeddings. Let’s break it down.
Embedding is the process of converting a word into a vector, which is just an ordered list of numbers with dimensions. Each number represents a dimension, and each dimension captures some aspect of a word’s meaning.
But here’s the kicker: we don’t manually define what each dimension means. There’s no lookup table saying, “dimension 48 is related to color.” The model figures all that out on its own during training by looking at how words show up in context across the traning material.
So even though individual numbers might not make much sense on their own, the whole vector ends up carrying enough information for the model to understand relationships between words - things like similarity, analogies, and usage patterns.
We call it a “high-dimensional space” because these vectors can have hundreds (potentially thousands) of dimensions. More dimensions usually means better representation and potentially more accurate results, like sharper sentiment analysis, but it also means more training time and more computational overhead. Also, cranking up the dimensionality too much can actually hurt performance in some cases by causing the model to overfit the training data, which makes it less effective on new, unseen text.
The goal of word embedding vectors is to help LLMs understand the meaning of words based on their context. Instead of treating words as isolated symbols, embeddings attempt to allow the model capture relationships and similarities between words. Word embeddings place each word into a high-dimensional vector space. In this space, similar words appear close together.
This enables models to do some surprisingly intuitive things, like analogies. For example: “Paris is to France as Tokyo is to Japan.”
This kind of analogy can be expressed using simple vector arithmetic. If you take the vector for Paris, subtract the vector for France, and add the vector for Japan, you land close to the vector for Tokyo. The model captures the relationship between capital cities and their countries. Pretty cool, right?
Further reading suggestion: Deep Learning for Natural Language Processing.
Bag of Words, Word2Vec, and Other Techniques
Embedding techniques laid the foundation for LLMs by showing how language could be represented numerically in a way that captures meaning. Some popular techniques that you might want to be aware of are Bag of Words and Word2Vec.
Bag of Words is a simple model that creates vector representations by counting how often each word appears in a document, ignoring grammar and word order. It is easy to implement but doesn’t capture word meaning or context.
Word2Vec is a more complex model developed at Google in 2013. It creates dense vector representations by training a neural network (recall that a neural network is just a bunch of weighted mathematical functions designed to find patterns in data) to predict a word based on its surrounding context, or predict the context from a given word. This results in embeddings where words with similar meanings are close together in the vector space.
There are other techniques, like GloVe, developed at Stanford in 2014, but we ain’t got time for all that here.
Simple LLMs might use Bag of Words or Word2Vec directly, often with pretrained embedding values. Many of today’s larger models use similar ideas, dense vector representations that capture meaning, but they typically start with random embeddings and learn everything from scratch. As training progresses, these embeddings are gradually adjusted, allowing the model to build its own custom vector space tuned to the data it sees.
Positional Embeddings
In principle, simple embedding vectors can serve as inputs to language models. These methods represent words without capturing their order or context, which limits their usefulness in understanding natural language.
Modern LLMs improve on this by adding on something called “position-aware embeddings.” The goal is to encode both the identity of a word and its position in the sequence in the embedding vector. This allows the model to distinguish between sentences which have the same words but different meanings due to word order.
There are two main types of position-aware embeddings:
- Absolute positional embeddings: Each position in the input sequence is assigned a fixed embedding vector. This vector is added to the word embedding before being passed into the model.
- Relative positional embeddings: Instead of assigning a fixed vector to each position, the model learns how the positions of words relate to one another. This can help the model generalize better across varying sequence lengths.
These techniques help the model understand the structure of a sentence and not just the individual words, which improves performance across many language tasks.
What is a Transformer?
A transformer is a kind of robot action figure that turns into a car or a truck or something. They were very popular in the 80s, 90s, and 2000s. But you’re probably here for the kind of transformer that powers LLMs, so fine, we’ll talk about that instead.
Most modern LLMs use something called the “transformer architecture” which was introduced in the 2017 paper Attention is All You Need by a group of researchers at Google. This paper is credited as a major turning point in the field of AI development and is now one of the most cited research papers of the 21st century. It’s pretty academic, but worth a read if you’re into that kind of thing.
Put simply, a transformer is an architecture pattern for building deep learning models.
Most modern LLMs you’ve heard of, like GPT, use the transformer architecture. People often use “transformers” and “LLMs” interchangeably, but that’s not entirely accurate. Transformers are a general-purpose architecture that can be used in many AI applications beyond LLMs, such as computer vision.
Okay, so a transformer is just an architecture pattern. I’ve been around the block. I know about architecture patterns. I lived through the “microservices” phase. So what does a transformer actually do?
At the core of a transformer are two main components: encoders and decoders. Encoders are used to read text, and understand input data. Decoders are used to generate output, like predicting the next word in a sentence. Both the encoder and decoder work with tokenization, embedding vectors, and all the foundational pieces we covered earlier.
What really sets transformers apart is a mechanism called attention, specifically self-attention. We’ll talk more about this in a bit, but briefly, attention is a technique that helps the model decide which parts of the input are most relevant to each word it is processing. It allows the model to consider relationships between words across the entire sentence, rather than just looking at them in order.
GPT, BERT, and Transformer Variations
I’ll give you one guess what the “T” in GPT and BERT stands for. It’s transformer. GPT stands for Generative Pre-trained Transformer, and BERT stands for Bidirectional Encoder Representations from Transformers.
These are two variants of the transformer architecture optimized for different tasks. As you might have guessed, GPT models focus on generating text by focusing on the decoder part of the transformer architecture. BERT, on the other hand, is designed for understanding tasks like sentiment analysis and text classification by focusing on the encoder part of the transformer. Many modern GPTs don’t even use encoders anymore - just focusing on the decoder part for next-word generation.
The first GPT was introduced by OpenAI in 2018. BERT was introduced by Google in 2018 as well. Since then, organizations like Meta have released GPT-style models such as LLaMA and BERT variants like RoBERTa.
Attention
Attention is a technique that tries to figure out the importance of each part of a sequence relative to the other parts. Basically, when I’m looking at one part of a sentence and trying to understand it, attention helps decide what other parts of the sentence I should also be looking at, and how important those parts are in shaping the meaning of the part I’m focused on.
Let’s take the sentence: “The rabbit sat on the floor because it was lazy.” When the model gets to the word “it”, it needs to decide what “it” refers to. The floor? Probably not. More likely, we can tell that it probably refers to the rabbit. But how do you know that? It’s because your brain is doing something similar to what the attention mechanism does. You use context clues and your own experience with language to figure out that “rabbit” makes more sense. In your mind, you give more importance (attention calls this “weight”) to “rabbit”. Attention works the same way, except it uses some sophisticated math to model that weighting. It helps the model assign more relevance to certain words, like “rabbit”, when trying to understand a sentence. This weighting helps the model understand relationships and dependencies between words, which contributes to how it interprets meaning.
Attention is the special sauce, so to speak, that makes LLMs work.
How Does Attention Work?
Attention is usually the hardest part for people to grasp about LLMs. It certainly was for me. It involves a non-trivial amount of math and some abstract thinking. If you’re reading this and thinking, “Attention is easy and fun!” - then congratulations. Now shut up, nerd.
Self-attention is the mechanism that allows each token in an input sequence to consider, or “attend to,” all the other tokens in the input sequence when computing its own representation. “Regular” attention can involve two separate input sequences, while self-attention is named that way because it operates within a single sequence of inputs.
In our earlier example, “The rabbit sat on the floor because it was lazy, “ each word gets to “look at” every other word to decide what matters. This is what allows the model to resolve references like “it” and understand how different parts of the sentence relate to one another.
In self-attention, the goal is to compute a context vector for each element in the input sequence. A context vector is essentially another enriched embedding vector that not only represents the current element and its positional information, but also considers the same information from every other element in the sequence.
It “considers” that information by looking at the embedding vectors of every other element in the input sequence. There are many ways to compute attention, and we’ll walk through them in the next sections. The output, the context vector, is a new, richer representation of the original token. It blends information from across the entire sequence, weighted by relevance. This is how the model understands nuance and relationships: every word becomes aware of what’s happening around it.
Further reading suggestion: Hands-On Large Language Models.
Scaled Dot Product Attention
Attention mechanisms differ in how they calculate the final context vector, and some methods get quite sophisticated.
One of the most popular methods is called scaled dot-product attention, introduced by the Attention is All You Need paper we mentioned earlier.
Here’s how scaled dot-product attention works, step by step:
-
First, every token’s original embedding vector is transformed into three different vectors called query, key, and value vectors. These are not pulled from thin air, they come from multiplying the original embedding vector by a set of learned weight matrices that are updated during training.
-
Next, we calculate the attention score between the query vector of the token we are focusing on and the key vectors of all the other tokens. To do this, for the token we are focusing on, we take the dot product of its query vector with the key vectors of every token in the sequence. Why dot product? The dot product measures how similar two vectors are. In attention, it tells us how well a token’s query matches another token’s key. A higher dot product means more relevance, so the model knows to focus more on that token.
-
These raw scores can sometimes be very large or very small, which can make training unstable. To fix this, we divide the scores by the square root of the dimension of the key vectors. This scaling keeps the values in a stable range - and is also why it’s called “scaled” dot product attention.
-
Then, we apply a Softmax function to these scaled scores to turn them into probabilities that sum to 1. This step turns the raw scores into attention weights, telling us how much focus each token deserves relative to the current one.
-
Finally, to get the context vector for the token we are processing, we multiply each token’s value vector by its corresponding attention weight. Tokens that scored higher contribute more to the final context vector. We sum up these weighted value vectors, resulting in a new, context-rich vector that combines information from across the entire sequence but weighted to emphasize the most relevant parts.
That context vector is what the model uses to represent the token with all the relevant information it needs to understand the sentence. In short, scaled dot-product attention lets the model be very smart about which parts of the sentence to focus on for each word, making it possible to understand complex language patterns and generate coherent, meaningful text.
Other Types of Attention
Scaled dot-product attention is the go-to in modern transformers, but it’s not the only one at the party when it comes to attention. For example, Bahdanau attention came even before scaled dot product attention, back in 2015, and was used in early translation models. It’s slower but helped pave the way for everything that came after. Multi-head attention runs multiple attention calculations in parallel so the model can focus on different parts of the sequence at once. Masked and causal attention are used in models like GPT to prevent the model from looking ahead when it’s predicting the next word.
You could spend years studying attention, and plenty of people have. But we’ve covered the essentials here: what attention is, why it matters, and how it helps LLMs like ChatGPT understand language.
Transformer Layer
At the core of the transformer encoder or decoder is something called a transformer layer. Each layer consists of a few main parts, typically a self-attention block and a feedforward block, often referred to together as transformer blocks.
Say our input is the sentence: “The rabbit ate the.” After tokenization and embedding, we get 4 vectors, one for each token (where there is a 1:1 mapping of tokens to word in this example). These vectors go into the first transformer block. The self-attention mechanism runs and produces 4 new vectors, called context vectors, which we covered earlier.
Next, the transformer block adds the original input vectors to these context vectors and applies normalization. Adding the original input helps preserve the starting information, so each transformer layer tweaks the data rather than changes it completely. Normalization stabilizes training by avoiding the problem of exploding or vanishing values, where numbers blow up to infinity or shrink to zero, derailing the model.
The result then flows into the feedforward block. Despite the name, it’s just a small neural network (a stack of mathematical functions applied independently to each token vector) with no interaction between tokens at this stage.
Inside the feedforward block, each vector gets expanded to a larger size, processed through activation functions, then shrunk back to its original size. These activation functions are constant mathematical functions, they don’t have any learnable parameters, but they introduce the nonlinear behavior that allows the model to learn complex patterns. Without them, the model could only perform simple linear transformations. For instance, a 768-dimensional vector might expand to 3072, get passed through one of these fixed nonlinear functions and then shrink back to 768. Common activation functions include ones like GELU, but we ain’t got time for all that here.
After the feedforward block, the input to that block is added back again, and normalization is applied one more time. So by now, the vectors have been enriched with context from attention, transformed nonlinearly, and stabilized withresidual connections and layer normalization.
That’s one transformer layer. GPT-style models stack a dozen or more of these, where the magic compounds. The same applies to BERT-style models, they stack multiple transformer layers, each built from self-attention and feedforward blocks, to deepen understanding. The main difference: encoders use bidirectional self-attention (seeing all tokens), while decoders use masked self-attention (to avoid looking ahead).
So in short: a transformer layer takes in vectors, from embeddings or a previous layer, runs attention, applies some mathematical magic, and sends out a fresh set of vectors.
Further reading suggestion: Build a Large Language Model (From Scratch)
Parameters
I know where you’re at right now, very far into this article, and it feels like all we’ve done is talk about vectors. You’re probably wondering: Where does this vector rainbow end? Don’t worry, we’re almost there. But first? Parameters.
A parameter (sometimes called a weight) is a trainable number inside the model, basically a knob the model can tweak to improve its performance. The more parameters a model has, the more complex patterns it can learn. GPT-2, for example, had a relatively modest 1.5 billion parameters. GPT-4? It’s rumored to have a staggering 1.76 trillion.
During training, the model adjusts these parameters to minimize its loss function, a number that tells it how far off its predictions are from the correct answers. Lower loss = better performance. That’s how the model learns.
These parameters live inside the transformer layers, but where exactly? In several key places.
In the self-attention components, the model has three distinct sets of parameter matrices: the query (Q) weight matrix, key (K) weight matrix, and value (V) weight matrix. These are learned during training and stay fixed during inference. When a token vector enters the layer, it’s multiplied separately by each matrix to produce the Q, K, and V vectors.
Next, the token vector flows into the feedforward block, which has its own set of parameters, typically two large weight matrices and some biases. First, the vector is multiplied by the first matrix, expanding its size. Then it passes through a nonlinear activation function (recall that these have no parameters), and finally through a second matrix that shrinks it back to the original size. These feedforward parameters are also learned during training and help the model transform each token’s representation in complex, nonlinear ways, but independently from the other tokens.
Aside from those two main blocks, parameters also show up in other spots, like in the normalization layers, where small learnable parameters help stabilize the data as it moves through layers. You’ll also find parameters in the embedding layers (where tokens are converted into vectors) and the output projection layer (where vectors are turned into token predictions).
Almost every step of the transformer architecture has some set of learnable weights. And during training, each of them is gradually adjusted to minimize loss and make the model better at predicting the next word, phrase, or sentence.
Output Generation
Alright, so we’ve got all these layers, all these parameters, and all these vectors flying around. But how does this actually turn into a prediction? How does ChatGPT, for example, decide what word comes next in a sentence?
After a token passes through the full stack of transformer layers, what comes out the other side is a single vector, the final context vector for that token. This vector is a condensed, context-rich representation of everything the model has learned about that token in relation to the others. It reflects the cumulative results of attention scores, feedforward transformations, layer normalizations, and parameter tuning.
But a vector isn’t a word, it’s just a list of numbers.
To turn that vector into an actual prediction, like the next word in a sentence, the model does one last thing: it multiplies the context vector by a big matrix called the output projection matrix. This matrix has one row for every token in the model’s vocabulary.
Why does this work? When the model multiplies the context vector by this matrix, it’s essentially checking: How similar is this context vector to each possible token? The result is a list of raw scores, one per token, showing how well each word fits the current context.
Here’s the trick: during training, the model learns to shape the context vector so that it ends up close (in vector space) to the row in the output projection matrix that corresponds to the correct next token. It does this by adjusting its parameters until the context vector lines up nicely with that token’s vector. The better the alignment, the higher the score after the multiplication.
But these scores aren’t probabilities yet, they’re just numbers. To turn them into actual probabilities, the model runs them through the Softmax function. Softmax takes all the scores, boosts the highest ones, shrinks the lower ones, and turns them into values between 0 and 1 that add up to 1. Now the model has a probability for every possible token.
Whichever token has the highest probability is usually selected, unless the model is using sampling or temperature (low temperature boosts the top probabilities, high temperature means more randomness) to add variety. That’s the final output. So, given something like “The rabbit ate the,” the model might assign high probabilities to “carrot,” “lettuce,” or “grass,” and very low ones to things like “motorcycle.”
And just like that, the next token is picked.
Training
But how does the model learn to make better predictions?
During training, the model is fed massive amounts of unlabeled data from the corpora. For every input sequence, it tries to predict the next token one at a time. After each guess, it compares its prediction to the actual next token and computes something called loss, a number that measures how wrong it was.
So how does the model measure “how wrong” it was? When predicting the next word, the model assigns a probability to every token in its vocabulary. Ideally, the correct token should have a probability close to 1, and all others near 0. The way the model measures the difference between its guess and the truth is called cross-entropy loss. This gives a single number that shows how far off the prediction is, zero means a perfect guess, and bigger numbers mean worse mistakes. Cross-entropy loss looks at the probability the model gave to the correct token. If that probability is high, the loss is small. If it’s low, the loss grows larger. In simple terms, the loss punishes the model more when it’s less confident about the right answer.
After calculating the loss, the model uses backpropagation to figure out how much each parameter contributed to the mistake. It does this by computing gradients, signals that show how changing each parameter would affect the loss. Parameters with bigger gradients need bigger adjustments. Using these gradients, the model slightly tweaks its parameters to reduce the loss, improving future predictions.
This cycle, predict, compare, correct, repeat, happens over and over again, billions of times across billions of tokens. Each iteration nudges the model toward slightly better performance. Over time, those tiny adjustments accumulate into a model that’s shockingly good at predicting what comes next.
In GPT-style models, this process is always left-to-right: predict the next token, then the one after that, and so on. In BERT-style models, it’s a bit different, the model tries to fill in missing words from both directions, but the idea is the same: guess a token, measure how wrong you were, and adjust the parameters to get it right next time.
That’s training in a nutshell, millions of micro-corrections that add up to a model that sounds like it actually understands language.
Foundational Models
A foundational model is a large, pre-trained general-purpose model trained on massive, diverse datasets. You can take a foundational model and fine-tune it to:
- Write code, like OpenAI Codex which was fine-tuned from GPT-3
- Analyze medical data, like Med-PaLM which was fine-tuned from Google’s PaLM
- Analyze biomedical or scientific literature, like SciBERT
Foundational models are like a blank canvas, some popular ones are:
- OpenAI: Creator of the GPT series (GPT-2, GPT-3, GPT-4), including Codex and ChatGPT.
- Anthropic: Creator of the Claude model family, focused on alignment and constitutional AI.
- Meta: Creator of the LLaMA (Large Language Model Meta AI) series, including open-weight models.
- Google DeepMind: Creator of the Gemini family, successor to PaLM, with multimodal capabilities.
- Mistral AI: Developer of lightweight, high-performance open-weight models like Mistral 7B and Mixtral.
Reinforcement and RLHF
So you’ve got yourself a foundational model, GPT, Claude, LLaMA, trained on a galaxy of internet text. It’s powerful, sure. But is it the best that it could be for you’re particular task? Probably not. That’s where fine-tuning comes in, specifically Reinforcement Learning from Human Feedback (RLHF).
Reinforcement learning is a training method where a model learns through trial and error: take an action, get a reward (or punishment), and adjust behavior to get better rewards next time.
Here’s how RLHF works in three steps:
- Supervised fine-tuning, First, the model is trained on human-written examples of “good” answers.
- Reward modeling, Then, humans compare several possible responses to a prompt and rank them best to worst. This allows the model to score the output.
- Reinforcement learning, Finally, the model generates responses, has them scored, and updates itself using reinforcement learning to maximize helpfulness and reduce weirdness.
RLHF essentially puts a human in the loop to guide the model’s behavior, someone who can say, “No, LLM, that wasn’t helpful. Please try again.”
There are other fine tuning techniques, of course, but we ain’t got time for that here.
Prompt Engineering
I’ll be honest, I assumed that the phrase “prompt engineering” was a scam. “You mean to tell me the trick to better AI is… phrasing the question differently?” Actually, yes. That does work (sometimes).
Now that we understand how these models work, it’s not that surprising. Remember back when we talked about attention? The model literally focuses on the parts of the input stream (i.e. your prompt) it thinks are important. So the more helpful context you include, tone, structure, examples, the easier it is for the model to lock onto what you actually want.
That’s the whole game: make the model’s job easier, and it gives you better answers.
In fact, the cooler kids are starting to call this context engineering, same idea as prompt engineering, but expanded. Instead of just rephrasing your prompt, you’re using everything in the model’s toolkit: long-term memory, retrieval, system instructions, and more.
Emergent Properties
Emergent properties are capabilities that weren’t explicitly trained, but somehow appear once the model is big enough. Like the model spontaneously learning how to do math, translate languages, or write poetry. Nobody explicitly taught it how to do those things, it just picked up those skills from the patterns in the training data.
As models scale, more parameters, deeper layers, more diverse training data, they hit a kind of complexity threshold. Once they cross that line, they stop just memorizing examples and start generalizing. Not because they were explicitly trained to reason or rhyme, but because it helped them minimize the loss function across billions of tokens.
Researchers started seeing this kind of emergent behavior in large models like GPT-3, where abilities like question answering, summarization, or code generation only showed up once the model got big enough.
Emergent properties are closely tied to zero-shot and one-shot learning, which basically describe how much help a model needs to perform a new task:
- In zero-shot, you give the model just the prompt, and it somehow figures out what to do.
- In one-shot, you show it one example before asking it to try.
It’s not magic, it’s math. But it feels magical when your autocomplete figures out how to summarize an academic paper after reading one sentence.
Hallucinations
You’ve seen it. I’ve seen it. We all know about it. But why does this happen?
Because LLMs aren’t search engines. They don’t fact-check. They don’t “know” things. They’re just extremely advanced guessers, predicting the next token based on patterns they’ve seen in their training data. If you ask about a fake law or a non-existent scientist, the model will happily invent one. Not because it wants to lie, but because, statistically, that’s what a plausible response looks like.
This is called a hallucination, when a language model confidently outputs something that sounds correct but is completely made up.
Hallucinations are a side effect of how LLMs are built. These models aren’t querying a database of facts, they’re generating responses by completing patterns. If the training data didn’t cover a topic accurately (or at all), the model will still try to fill in the blanks with something that looks right.
Some newer models try to reduce hallucinations by connecting the LLM to external tools, like search engines or fact-checking APIs, like:
- OpenAI’s ChatGPT with browsing uses Bing to fetch real-time web results.
- Google’s Gemini integrates search and fact-checking pipelines under the hood.
But unless you’re using one of those fancier setups, always take a model’s answers with a grain of salt.
Copyright Considerations
Here’s the billion-dollar question: if LLMs are trained on human-created content, who owns what they generate?
These models learn from mountains of text, code, music, and art, much of it copyrighted. That raises some thorny legal questions. Is AI-generated output original, or a remix? What is “fair use” for LLMs? If models were trained on open-source code, do the outputs inherit the license?
Right now, the law hasn’t caught up. Lawsuits are in progress, including cases against OpenAI (NYT v. OpenAI) and GitHub Copilot (Doe v. GitHub), and regulators are still figuring it out.
GPU Mania
Training and running LLMs involves performing billions (or trillions) of matrix operations. GPUs (Graphics Processing Units) are optimized for exactly that kind of parallel math. Unlike CPUs, which handle tasks sequentially, GPUs can process thousands of operations simultaneously, making them ideal for deep learning workloads.
Without GPUs, training a model like GPT-4 would take years, or just plain fail.
NVIDIA dominates the space, powering most major LLMs today (NVIDIA Blog) - so now you know why NVIDIA’s stock went bananas in the 2020s.
Conclusion
Well, we made it.
We’ve gone from “What even is an LLM?” to deep-diving into embeddings, attention, training loops, fine-tuning, hallucinations, and legal gray zones.
So what’s the big takeaway? LLMs aren’t magic. They’re just really good at spotting patterns in massive amounts of text, powered by math, scale, and a bit of human feedback.
And now that you know a little about how they work, you’ll be in a much better position to spot the real deal versus the phonies.
