Training data
- Loss (cross-entropy)
- The average −ln(probability the model gave the correct next word). Lower is better; 0 means perfect prediction.
- Perplexity = eloss
- How many words the model is effectively torn between. 1 = sure and right; higher means more confused (e.g. 5 ≈ guessing among ~5 words).
- Log-likelihood = −loss
- The average log-probability of the correct words. It's ≤ 0; closer to 0 (less negative) is better.
The network — watch the weights change as it learns
Blue = positive weight, orange = negative; thicker = stronger. Output dots light up with the predicted next word. Pause training anytime to inspect.
Predict the next word
While training, this plays through real examples: the 4 input words change, the bars show the model's guesses, and the true next word is marked in red (green when the model gets it right). Pause anytime to try your own words.
🏆 Word Predictor Challenge
Turn on Competition mode to train on a fixed, shared corpus. It's split into train / held-out test, and your model is scored on how well it predicts the next word it has never seen — rewarding accuracy and a small model. Uploading your own text is disabled while competing.
Word embedding map
Each word is a point. As the model learns, words it uses the same way move together.