Learning Hub/Machine Learning/Session 5

Session 5 of 5 · 75 minutes

Can We Trust the Model?

Spot overfitting, think about fairness, and finish with a classifier project you can present.

Goals

By the end of this session you can:

  • Explain overfitting by comparing training accuracy with testing accuracy
  • Give two ways that data can make a model unfair
  • Complete and present a small classification project
  • Describe how machine learning could help solve a problem in your community

Words to know

overfitting
When a model memorizes its training data and does badly on new data.
bias (in data)
When the data leaves out or misrepresents some people or situations, so the model is unfair to them.
decision tree
A model that asks a series of yes/no questions to reach an answer.
fairness
Treating people equitably. A fair model works well for everyone it will affect.

The lesson

Step 1: A model that asks questions

A decision tree reaches an answer by asking a series of yes/no questions about the features: “Is the alcohol level above 13? If yes, is the color intensity above 4?…” Each question splits the data into smaller groups.

We will use another dataset that ships with scikit-learn: 178 wines, each described by 13 measurements and labeled with one of three growers.

Python
from sklearn.datasets import load_wine
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

wine = load_wine()
X_train, X_test, y_train, y_test = train_test_split(
    wine.data, wine.target, test_size=0.3, random_state=0
)
print(len(wine.data), "wines,", len(wine.feature_names), "features")
Output
178 wines, 13 features

Step 2: When a model memorizes

A tree can keep asking questions until it has a separate answer for every training example. That is overfitting. Let’s compare training accuracy with testing accuracy as we let the tree grow deeper.

Python
for depth in [1, 2, 3, None]:
    tree = DecisionTreeClassifier(max_depth=depth, random_state=0)
    tree.fit(X_train, y_train)
    train_acc = accuracy_score(y_train, tree.predict(X_train))
    test_acc = accuracy_score(y_test, tree.predict(X_test))
    print(depth, round(train_acc, 3), round(test_acc, 3))
Output
1 0.653 0.63
2 0.952 0.87
3 0.992 0.944
None 1.0 0.944

The columns are: depth limit, training accuracy, testing accuracy. (None means no limit.)

  • At depth 1, the tree is too simple. It is wrong on both sets (underfitting).
  • With no limit it scores a perfect 1.0 on training data. But on new data it does no better than depth 3.

The gap between training and testing accuracy is the warning sign. Always look at testing accuracy, never training accuracy alone.

Step 3: Whose data is it?

A model learns from the data we give it. If the data is unbalanced, the model will be too.

  • A hiring model trained on past hires may learn to prefer the kinds of people who were hired before, even if they were not the best.
  • A speech tool trained mostly on one accent may misunderstand others.
  • A health model trained on one population may not work for another.

Nobody has to mean harm for this to happen. It comes from who is in the data and who is missing.

Step 4: Questions for a fairer model

Before you trust a model, ask:

  1. Who is in the data? And who is left out?
  2. How does it perform for different groups? Don’t just look at overall accuracy.
  3. What happens when it’s wrong? Is a mistake annoying, or does it harm someone?
  4. Who decides? For important choices (school, jobs, health) a person should review the model’s output.
  5. Can people question the result? A fair system has a way to say “this is wrong”.

Think about it: Machine learning can also do a lot of good: reading X-rays in places with few doctors, forecasting floods, helping farmers spot crop disease. What problem in your community could use data?

Step 5: Your capstone

You now know the whole path: ask a question, collect data, look at it, train a model, test it fairly, and reflect on whether you can trust it. For the capstone project, choose a dataset and a depth, and prepare a one-minute talk with four parts:

  1. The problem you tried to solve.
  2. Your best model and its testing accuracy.
  3. One thing that could go wrong.
  4. One real-world use you’d be excited about, or careful with.

Good presentations are honest about both what works and what doesn’t. That honesty is what makes a data scientist trustworthy.

Exercise

Spot the problem

  1. A model for hiring was trained on past hires, and most past hires were men. What could go wrong?
  2. A voice assistant was trained mostly on one accent. Who might it fail?
  3. A photo model gets 99% accuracy on training photos but 60% on new ones. What is happening?
  4. For each case, write one idea to make it better.
Need a hint?

Ask: who is in the data? Who is missing? What does the model do for the people who are missing?

Small project

Capstone: build and present a classifier

Pick a dataset, train a classifier, and give a one-minute presentation on how well it works, where it fails, and why that matters.

  1. Choose a dataset from scikit-learn: load_wine(), load_breast_cancer() or load_iris(), or one from your teacher.
  2. Split into training and testing sets. Train a DecisionTreeClassifier.
  3. Report training accuracy and testing accuracy. Is there a gap?
  4. Try at least three values of max_depth and record the results in a table.
  5. Prepare a one-minute talk: the problem, your best model and its accuracy, one thing that could go wrong, and one real-world use you'd be careful with.

Stretch it: Connect it to your community: describe a real problem near you (water, farming, health, transport) and what data you would need to collect to try machine learning on it.

Quiz

Check your understanding

Pick one answer for each question. Your score appears right in the page; nothing is sent anywhere.

  1. Question 1A model scores 100% on its training data but 62% on testing data. This is a sign of…
  2. Question 2Which situation is most likely to produce an unfair model?
  3. Question 3What does a decision tree do?
  4. Question 4Limiting a tree's max_depth usually…
  5. Question 5Before using a model that affects people (for example, grading or loans), what should you do?