Logo
All posts

The curious case of Attention in LLMs

When you type "apple" into ChatGPT, how does it distinguish between fruits and phones?

Gaurav Sen/6 min read
One word, two things. "Apple" is a fruit in one sentence and a phone in the next.
One word, two things. "Apple" is a fruit in one sentence and a phone in the next.

"Selling apples in the market" and "selling Apple phones" start with the same vector for apple. The mapping is one-to-one.

Two sentences about two different apples. Both hand the model the same starting vector.
Two sentences about two different apples. Both hand the model the same starting vector.

So how does the LLM understand ambiguous terms like apple or match (cricket match and a stick of wood)?

Engineers in our AI engineering cohorts asked this 54 times. Almost half of them were still unsure after the first answer.

So let's take the questions in the order they came.

Why does "apple" start as the same vector every time?

A token is the unit of text a model reads. Every token has an ID, and every ID maps to exactly one vector.

Look at that vector alone and you can't tell which apple it is. The model can't either.

Then how does the model know which apple I mean?

The words around it tell the model.

In "Apple 17 Pro", attention lets "Pro" push the apple vector towards phones. In "apple pie", the word "pie" pushes the same vector towards food.

One starting vector for apple. "Pie" moves it towards the fruit, and "Pro" moves it towards the phone.
One starting vector for apple. "Pie" moves it towards the fruit, and "Pro" moves it towards the phone.

"Match" goes the same way. "Cricket" pushes it towards a game, and "matchbox" pushes it towards a stick of wood.

Attention is the operation that rewrites each token's vector as a weighted sum of the vectors around it, weighted by how relevant each of those tokens is to it.

So two different vectors get called "the embedding". The starting vector is the same every time. The vector after attention changes with every sentence.

What if "apple" appears twice in one sentence?

Try "apple pie is the apple of my eye".

The two apples stand at positions 1 and 5. They start as the same vector, and position is the first thing that tells them apart.
The two apples stand at positions 1 and 5. They start as the same vector, and position is the first thing that tells them apart.

Both apples start as the same vector, like A = 3 and B = 3. Their positions differ, so the position encoding moves them apart a little.

Attention does the rest. "Pie" pushes the first apple towards food, and "eye" pushes the second towards a feeling.

Where do the starting vectors come from?

Random numbers. You don't decide them.

The model is asked to predict the next token. When the prediction is bad, backpropagation changes the vectors, and when it is good, they stay.

This makes the vectors parameters of the model, like every other weight.

Where are the vectors stored? Is there a vector database inside the model?

There is no vector database inside the model.

The starting vectors sit in one table inside the weights, with one row per token. GPT-2 has 50,257 rows of 768 numbers each, and the model looks up a row by token ID.

The vectors after attention are not stored anywhere. They are computed for your request and thrown away when the answer is done.

What does each of the 768 numbers mean?

We don't know.

When I draw an axis and label it "happiness", I am lying to you a little. Emotion and intensity could be mixed into one dimension, because nobody labelled them.

Think of the dimensions as columns in an Excel sheet. Every token has the same columns, and training fills in the values.

Is "capital of" stored somewhere?

No vector holds that relationship.

Take the vector for Bengaluru, subtract Karnataka and add Tamil Nadu. You land near Chennai!

Four arrows from a capital to its state, drawn on toy coordinates. All four point roughly the same way.
Four arrows from a capital to its state, drawn on toy coordinates. All four point roughly the same way.

Nobody told the model to arrange cities this way. We only punished wrong predictions of the next word, and the vectors moved until this pattern appeared. It is observed, not forced.

Why 768 dimensions? Why not more?

The size is fixed by the people who design the model. GPT-2 used 768, and larger models use several thousand.

More dimensions carry more detail, and they cost more, because every attention step multiplies matrices of that width.

For RAG, don't guess. Take 5% to 10% of your documents, run real queries against two embedding sizes, and compare accuracy with the bill.

Summary

  1. A token always starts as the same vector, whatever the sentence.
  2. Attention rewrites that vector using the tokens around it. This is how the fruit and the phone separate.
  3. Starting vectors are rows of a table inside the model's weights. They begin as random numbers and are learned in training.
  4. Vectors computed during a request are thrown away after the answer.
  5. No dimension has a known meaning, and no vector stores a relationship like "capital of".
  6. A larger embedding size costs more compute at every attention step.
WhatsApp