🎬 Movie Rating Predictor — real IMDb data

Can a network guess a film's IMDb rating before the reviews are in? Pick your inputs — year, duration, popularity, genre — and two engineered features: the director's and lead actor's reputation (their average rating on other films). Shape the network and train (max 50 epochs). Real ratings are noisy, so a good fit is hard-won. Score = test MSE × (1 + 0.4% per parameter), up to 3 attempts.

Data: IMDb Indian Movies (Kaggle) — top 3,500 most-voted rated films. Reputation features are built from the training split only (leave-one-out), so they're fair, not leaky.

0

FEATURES

What predicts a rating? Not everything does. Click to toggle. 3 selected.

2 hidden layers

Line blue = positive weight, orange = negative; thickness = strength. Watch them shift as it trains.

OUTPUT

Train MSE
—
Test MSE
—
R² (fit)
—
🏆 Your Score lower is better
—
test MSE × (1 + 0.4% per parameter)
Test error
Size penalty

• train   • test   — perfect fit

MSE per epoch — train / test

🍿 Rate a real movie

—
—
model predicts
—
actual IMDb
train the model to predict

🧪 Things to try

One-hot vs reputation

Toggle a few “Directed by …” one-hot inputs, then swap them for Director reputation. One-hot only helps that one director and can't generalize; the single reputation number captures every director. That's why we encode high-cardinality categories instead of one-hot-ing them.

Popularity ≠ quality

Popularity (vote count) nudges the prediction, but blockbusters aren't always well-rated. See how far it gets you alone.

Genre is weak

Genre flags barely move the needle — a hint that some inputs are near-useless and just cost parameters. Drop them to lower your Score.

Real data is humbling

Even your best model leaves lots of scatter (R² well under 100%). You can't fully predict taste from metadata — recognising that ceiling is the lesson.