Training data

Epoch
0
Loss
—
Perplexity
—
Log-likelihood
—
Loss (cross-entropy) per epoch
Loss (cross-entropy)
The average −ln(probability the model gave the correct next word). Lower is better; 0 means perfect prediction.
Perplexity = eloss
How many words the model is effectively torn between. 1 = sure and right; higher means more confused (e.g. 5 ≈ guessing among ~5 words).
Log-likelihood = −loss
The average log-probability of the correct words. It's ≤ 0; closer to 0 (less negative) is better.

The network — watch the weights change as it learns

Blue = positive weight, orange = negative; thicker = stronger. Output dots light up with the predicted next word. Pause training anytime to inspect.

Predict the next word

While training, this plays through real examples: the 4 input words change, the bars show the model's guesses, and the true next word is marked in red (green when the model gets it right). Pause anytime to try your own words.

Your 4 words
→
Embeddings
→
Model's next word

🏆 Word Predictor Challenge

Turn on Competition mode to train on a fixed, shared corpus. It's split into train / held-out test, and your model is scored on how well it predicts the next word it has never seen — rewarding accuracy and a small model. Uploading your own text is disabled while competing.

Word embedding map

Each word is a point. As the model learns, words it uses the same way move together.