Session 3 of 5 · 60 minutes
Predicting Numbers
Train a first model with scikit-learn that predicts a number from another number.
Goals
By the end of this session you can:
- Describe a regression problem: predicting a number
- Fit a linear regression model with scikit-learn
- Use the model to predict for new inputs
- Split data into training and testing sets and measure the error
Words to know
- regression
- A machine learning task where the answer is a number, like a price or a score.
- scikit-learn
- A popular Python library with ready-made machine learning tools.
- fit
- To train a model on data.
- predict
- To use a trained model on new input.
- error
- How far a prediction is from the real answer.
The lesson
Step 1: Predicting a number
Last time we saw that more study time goes with higher scores. Can we turn that pattern into a prediction?
When the answer we want is a number (a score, a price, a temperature), the task is called regression. The simplest approach is to draw the best straight line through the dots and read predictions from it. The computer finds that line for us.
Step 2: Fit a model
scikit-learn provides ready-made models. We will use LinearRegression.
import pandas as pd
from sklearn.linear_model import LinearRegression
data = pd.DataFrame({
"hours": [1, 2, 3, 4, 5, 6, 7, 8],
"score": [52, 55, 61, 64, 70, 74, 80, 85],
})
X = data[["hours"]] # input (a table, so two brackets)
y = data["score"] # the answer we want to predict
model = LinearRegression()
model.fit(X, y)Two names are used everywhere in machine learning: X for the input features and y for the answers (the labels). fit() is the training step.
Step 3: Look at the line and make predictions
A line has a slope and an intercept. The model stores both.
print("slope:", round(model.coef_[0], 2))
print("intercept:", round(model.intercept_, 2))slope: 4.77
intercept: 46.14So the model’s rule is: score ≈ 4.77 × hours + 46.14. Each extra hour of study adds about 4.77 points.
Now ask it for predictions on new inputs:
new_students = pd.DataFrame({"hours": [4.5, 10]})
print(model.predict(new_students).round(1))[67.6 93.9]The prediction for 4.5 hours is believable. The prediction for 10 hours is outside anything the model saw, so it is more of a guess.
Step 4: A fair test
A model can look perfect on the data it learned from. To be fair we hide some data, train on the rest, then see how it does on the hidden part.
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=1
)
model = LinearRegression().fit(X_train, y_train)
predictions = model.predict(X_test)
print("test hours:", list(X_test["hours"]))
print("actual:", list(y_test))
print("predicted:", predictions.round(1))
print("MAE:", round(mean_absolute_error(y_test, predictions), 2))test hours: [8, 3]
actual: [85, 61]
predicted: [83.9 60.3]
MAE: 0.9test_size=0.25keeps 25% of the rows for testing.random_state=1makes the shuffle the same every time so we can compare results.- MAE (mean absolute error) is the average size of the mistakes: our predictions are about 0.9 points off.
Your numbers may differ a little if your library version is different.
Step 5: What this model can and can’t do
Our line is clean because the data was made up to be nearly a straight line. Real data is messier, and a straight line may not fit well.
Remember these cautions:
- Only eight students. More data would make the model more trustworthy.
- A line goes on forever. At 20 hours the model would predict a score over 100, which is impossible.
- Other things matter. Sleep, topic difficulty and teaching quality also affect scores, but the model only knows about hours.
A good data scientist always reports how well the model works, where it works, and where it doesn’t.
Exercise
Predict with your own data
- Make a DataFrame with a column of
xvalues (for example, hours of practice) and a column ofyvalues (for example, free-throws made) with at least 8 rows. - Fit a
LinearRegressionmodel withxas input andyas the answer. - Print the slope and intercept and say in words what the slope means.
- Predict
yfor a newxvalue you didn't use in training.
Need a hint?
Remember the input must be two-dimensional: data[["x"]] with double brackets.
Small project
Score predictor
Build a model that predicts a student's test score from the hours they studied, and test how accurate it is.
- Load the study-time data from Session 2.
- Split it with
train_test_split(keep 25% for testing). - Fit a linear regression on the training part.
- Predict the testing part and compute the mean absolute error.
- Ask your model: what score do you predict for 4.5 hours? For 10 hours?
- Write two sentences: Is the answer for 10 hours reliable? Why or why not?
Stretch it: Plot the data points and draw the line the model learned over them.
Quiz
Check your understanding
Pick one answer for each question. Your score appears right in the page; nothing is sent anywhere.
Feedback on this lesson
Teachers and students: tell us what worked and what didn't. We read this to improve the curriculum.
Received
Thank you for the feedback.
The Education team reviews lesson feedback to improve the curriculum.
Preview mode: this form is not connected to a destination yet, so nothing was sent. Set its web address in src/data/config.json.