The Transformer on this page

This is one transformer block — the same machinery behind GPT, shrunk so you can watch every part. It reads your 4 words and predicts the next. Data flows top ā†’ bottom, and the shapes on the right update with your settings:

residual (skip) Your 4 words Token + Positional embeddings 4 Ɨ 8 → Q Ā· K Ā· V Multi-Head Self-Attention Ā· 1 head scaled dot-product → softmax → mix Values each head d→8 āŠ• Add & carry forward Feed-forward network d → 12 → vocab Softmax Next-word probabilities 23 words

Simplified from a full transformer. A production model (e.g. GPT) stacks many of these blocks; each block adds Layer Norm and a separate feed-forward sub-layer, uses full self-attention (every word attends to every earlier word) with causal masking, and a far larger vocabulary. Here we use one block, last-word attention, and fold the feed-forward into the output head — the ideas are identical, just smaller.

Training data

Epoch
0
Loss
—
Perplexity
—
Log-likelihood
—
Loss (cross-entropy) per epoch

Predict the next word

Your 4 words
→
Embeddings
→
Model's next word

Attention

When predicting, the last word looks back at all 4 words. Thicker green = more attention. Watch it change as the model learns (and when you change the words).

✨ Use it: generate text

A transformer's real job: predict a word, add it, slide the 4-word window, and repeat — exactly how ChatGPT writes. It starts from your 4 words above. Train first, then Generate.

🧪 Parameters & things to try

Attention heads

1 vs 2 — with two heads, each can focus on different words. Train, then compare the green vs. orange attention patterns.

Embedding dims

Size of each word vector (and the block's width). Bigger = more capacity, harder to picture. Changing it rebuilds the model.

Head neurons

Width of the feed-forward output head (d → this → vocab). More can fit trickier patterns.

Learning rate

0.03 is steady; 0.1 is faster but can wobble. Watch the loss curve.

Creativity (temperature)

The generation dial: 0 always picks the top word (repetitive); higher is more varied but riskier — the same knob real LLMs expose.

Upload .txt

Train on your own text, then Generate from your 4 seed words.

  1. Do heads specialize? Set Attention heads to 2, Train, and compare the two colored attention patterns.
  2. Greedy vs. creative: After training, Generate at temperature 0, then near 1 — repetition vs. variety.
  3. Train longer: Generation gets more sensible as the loss drops — keep training and re-Generate.
  4. Your text: Upload a short .txt and watch it continue your seed words.