Learning Hub/Machine Learning/Session 3

Session 3 of 5 · 60 minutes

Predicting Numbers

Train a first model with scikit-learn that predicts a number from another number.

Goals

By the end of this session you can:

  • Describe a regression problem: predicting a number
  • Fit a linear regression model with scikit-learn
  • Use the model to predict for new inputs
  • Split data into training and testing sets and measure the error

Words to know

regression
A machine learning task where the answer is a number, like a price or a score.
scikit-learn
A popular Python library with ready-made machine learning tools.
fit
To train a model on data.
predict
To use a trained model on new input.
error
How far a prediction is from the real answer.

The lesson

Step 1: Predicting a number

Last time we saw that more study time goes with higher scores. Can we turn that pattern into a prediction?

When the answer we want is a number (a score, a price, a temperature), the task is called regression. The simplest approach is to draw the best straight line through the dots and read predictions from it. The computer finds that line for us.

Step 2: Fit a model

scikit-learn provides ready-made models. We will use LinearRegression.

Python
import pandas as pd
from sklearn.linear_model import LinearRegression

data = pd.DataFrame({
    "hours": [1, 2, 3, 4, 5, 6, 7, 8],
    "score": [52, 55, 61, 64, 70, 74, 80, 85],
})

X = data[["hours"]]   # input (a table, so two brackets)
y = data["score"]     # the answer we want to predict

model = LinearRegression()
model.fit(X, y)

Two names are used everywhere in machine learning: X for the input features and y for the answers (the labels). fit() is the training step.

Step 3: Look at the line and make predictions

A line has a slope and an intercept. The model stores both.

Python
print("slope:", round(model.coef_[0], 2))
print("intercept:", round(model.intercept_, 2))
Output
slope: 4.77
intercept: 46.14

So the model’s rule is: score ≈ 4.77 × hours + 46.14. Each extra hour of study adds about 4.77 points.

Now ask it for predictions on new inputs:

Python
new_students = pd.DataFrame({"hours": [4.5, 10]})
print(model.predict(new_students).round(1))
Output
[67.6 93.9]

The prediction for 4.5 hours is believable. The prediction for 10 hours is outside anything the model saw, so it is more of a guess.

Step 4: A fair test

A model can look perfect on the data it learned from. To be fair we hide some data, train on the rest, then see how it does on the hidden part.

Python
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=1
)

model = LinearRegression().fit(X_train, y_train)
predictions = model.predict(X_test)

print("test hours:", list(X_test["hours"]))
print("actual:", list(y_test))
print("predicted:", predictions.round(1))
print("MAE:", round(mean_absolute_error(y_test, predictions), 2))
Output
test hours: [8, 3]
actual: [85, 61]
predicted: [83.9 60.3]
MAE: 0.9
  • test_size=0.25 keeps 25% of the rows for testing.
  • random_state=1 makes the shuffle the same every time so we can compare results.
  • MAE (mean absolute error) is the average size of the mistakes: our predictions are about 0.9 points off.

Your numbers may differ a little if your library version is different.

Step 5: What this model can and can’t do

Our line is clean because the data was made up to be nearly a straight line. Real data is messier, and a straight line may not fit well.

Remember these cautions:

  1. Only eight students. More data would make the model more trustworthy.
  2. A line goes on forever. At 20 hours the model would predict a score over 100, which is impossible.
  3. Other things matter. Sleep, topic difficulty and teaching quality also affect scores, but the model only knows about hours.

A good data scientist always reports how well the model works, where it works, and where it doesn’t.

Exercise

Predict with your own data

  1. Make a DataFrame with a column of x values (for example, hours of practice) and a column of y values (for example, free-throws made) with at least 8 rows.
  2. Fit a LinearRegression model with x as input and y as the answer.
  3. Print the slope and intercept and say in words what the slope means.
  4. Predict y for a new x value you didn't use in training.
Need a hint?

Remember the input must be two-dimensional: data[["x"]] with double brackets.

Small project

Score predictor

Build a model that predicts a student's test score from the hours they studied, and test how accurate it is.

  1. Load the study-time data from Session 2.
  2. Split it with train_test_split (keep 25% for testing).
  3. Fit a linear regression on the training part.
  4. Predict the testing part and compute the mean absolute error.
  5. Ask your model: what score do you predict for 4.5 hours? For 10 hours?
  6. Write two sentences: Is the answer for 10 hours reliable? Why or why not?

Stretch it: Plot the data points and draw the line the model learned over them.

Quiz

Check your understanding

Pick one answer for each question. Your score appears right in the page; nothing is sent anywhere.

  1. Question 1Which is a regression problem?
  2. Question 2What does model.fit(X, y) do?
  3. Question 3In score = 4.77 × hours + 46.14 what does 4.77 tell us?
  4. Question 4Why do we test the model on data it was not trained on?
  5. Question 5A model trained on hours between 1 and 8 predicts for 40 hours. How much should we trust it?