What is a Context Window? Unlocking LLM Secrets
Sintesi
Introduction to Context Windows
In the context of large language models, what is a context window? Well, it's the equivalent of the model's working memory. It determines how long a conversation the LLM can carry out without forgetting details from earlier in the exchange. Allow me to illustrate this using the scientifically recognized IBU scale—that's "International Blah Units." Here, "blah" represents me sending a prompt to an LLM chatbot. The chatbot then returns with a response—blah—and we continue the conversation.
Illustrating Context Window Size
I say something else, and it responds back to me. Blah blah blah blah—International Blah Units. This box here represents the context window, and in this case, the entire conversation fits within it. That means that when the LLM generated this response here—this blah—it had within its working memory my prompts to the model, here and here. It also had the other response that the model had returned to me in order to build this new response. All good. Now, let's consider a longer conversation.
Context Window Limitations
More blahs. I send my prompt—blah. It then sends me a response. Now we go back and forth with more conversation. I say something, it responds to that. I say one more thing, and it responds again. Now we have a longer conversation to deal with, and it turns out that this conversation thread is longer than the context window of the model. That means that the blahs from earlier in the conversation are no longer available to the model. It has no memory of them when generating new responses.
Tokens as Measurement Unit
The LLM can do its best to infer what came earlier by looking at the conversation that is within its context window. But now the LLM is making educated guesses, and that can result in some wild hallucinations. Understanding how the context window works is essential to getting the most out of LLMs. Let's get into a bit more detail about that now. My producer is telling me that context window size is, in fact, not measured in IBUs and that I made that up. We measure context windows in something called tokens.
Understanding Tokenization
Let's describe tokenization. Let's get into context length and size, and we're going to talk about the challenges of long context windows. To start, what is a token? For us humans, the smallest unit of information that we use to represent language is a single character—something like a letter, a number, or a punctuation mark. But the smallest unit of language that AI models use is called a token. A token can represent a character as well, but it might also be part of a word, a whole word, or even a short multi-word phrase. For example, let's consider the different roles played by the letter "A."
Token Examples
I'm going to write some sentences, and we're going to tokenize them. Let's start with "Martin drove a car." Here, "a" is an entire word, and it will be represented by a distinct token. Now, what if we try a different sentence: "Martin is amoral." Not sure why we would say that, but look—in this case, "a" is not a word on its own, but it's an addition to "moral" that significantly changes the meaning of that word. Here, "amoral" would be represented by two distinct tokens: one for "a" and another for "moral." One more: "Martin loves his cat." Now, the "a" in "cat" is simply a letter in a word. It carries no semantic meaning by itself and would therefore not be a distinct token. The token here is just "cat."
Tokenizer Functionality
The tool that converts language to tokens has a name—it's called a tokenizer. Different tokenizers might tokenize the same passage of writing differently. But as a general rule of thumb, a regular word in the English language is represented by about 1.5 tokens by the tokenizer. So, a hundred words might result in 150 tokens.
Self-Attention Mechanism
Context windows consist of tokens, but how many tokens are we talking about? To answer that, we need to understand how LLMs process tokens in a context window. Transformer models use something called the self-attention mechanism. The self-attention mechanism is used to calculate the relationships and dependencies between different parts of an input—words at the beginning and at the end of a paragraph, for example. The self-attention mechanism computes vectors of weights, in which each weight represents how relevant that token is to the other tokens in the sequence. The size of the context window determines the maximum number of tokens that the model can pay attention to at any one time. Context window size has been rapidly increasing. The first LLMs that I used had context windows of around 2,000 tokens. The IBM Granite 3 model today has a context window of 128,000 tokens, and other models have even larger context windows. But it almost seems like overkill, doesn't it? I would have to be conversing with a chatbot all day to fill a 128K token window. Well, that's not necessarily true, because there can be a lot of things taking up space within a model's context window. Let's take a look at what some of those things could be. One of them is the user input—the "blah" that I sent into the model. Of course, we also have the model responses as well—the "blahs" that it was sending back. But a context window may also contain all sorts of other things as well. Most models provide what is called a system prompt into the context window. This is often hidden from the user. But it conditions the behavior of the model, telling it what it can and cannot do. A user may also choose to attach some documents into their context window, or they might put in some source code as well. That can be used by the LLM to refer to in its responses. Supplementary information drawn from external data sources—for retrieval-augmented generation (RAG)—might also be stored within the context window during inference. A few long documents and some snippets of source code can quickly fill up a context window. Is a bigger context window always better? Well, larger context windows do present some challenges as well. What sort of challenges? I think the most obvious one would have to be compute. The compute requirements scale quadratically with the length of a sequence. What does that mean? As the number of input tokens doubles, the model needs four times as much processing power to handle it. Remember, as the model predicts the next token in a sequence, it computes the relationships between that token and every single preceding token in the sequence. As context length increases, more and more computation is required. Long context windows can also negatively affect performance, specifically the performance of the model. Both people and LLMs can be overwhelmed by an abundance of extra detail. They can also get lazy and take all sorts of cognitive shortcuts. A 2023 paper found that models perform best when relevant information is towards the beginning or towards the end of the input context, and that performance degrades when the model must carefully consider information that is in the middle of a long context. Finally, we also have to be concerned with a number of safety challenges. A longer context window might have the unintended effect of presenting a larger attack surface for adversarial prompts. A long context length can increase a model's vulnerability to jailbreaking, where malicious content is embedded deep within the input, making it harder for the model's safety mechanisms to detect and filter out harmful instructions. No matter how you measure it—either with IBUs or, more accurately, tokens—selecting the appropriate number of tokens for a context window involves balancing the need to supply ample information for the model's self-attention mechanism with the increasing demands and performance issues those additional tokens may bring.