Learning Hub/Machine Learning/Session 2

Session 2 of 5 · 60 minutes

Looking at Data

Use Python and pandas in Google Colab to load a small dataset, summarize it and spot a pattern.

Goals

By the end of this session you can:

  • Open a notebook in Google Colab and run Python cells
  • Create a pandas DataFrame and read its rows and columns
  • Summarize data with head(), describe() and mean()
  • Draw a simple chart and describe the pattern you see

Words to know

Google Colab
A free website where you run Python in notebook cells. No installation needed.
pandas
A Python library for working with tables of data.
DataFrame
A table in pandas, with rows and named columns.
feature
A column of facts used to make a prediction.
correlation
How strongly two columns move together, from -1 to 1.

The lesson

Step 1: Meet Google Colab

Google Colab (colab.research.google.com) runs Python in your browser. Sign in with a Google account and choose New notebook. A notebook is made of cells. Click in a cell, type code and press Shift + Enter to run it.

Python
print("Hello from Colab")
Output
Hello from Colab

Colab already has the libraries we need, so we can start right away.

Step 2: A table in pandas

We will use a tiny made-up dataset about students: the hours each one studied for a test and the score they got.

Python
import pandas as pd

data = pd.DataFrame({
    "hours": [1, 2, 3, 4, 5, 6, 7, 8],
    "score": [52, 55, 61, 64, 70, 74, 80, 85],
})
print(data.head())
Output
   hours  score
0      1     52
1      2     55
2      3     61
3      4     64
4      5     70

Each row is one student. Each column is one kind of fact. The numbers on the left (0, 1, 2…) are the row numbers pandas adds for you.

Step 3: Summarizing

describe() calculates a summary for every number column.

Python
print(data.describe().round(2))
Output
       hours  score
count   8.00   8.00
mean    4.50  67.62
std     2.45  11.72
min     1.00  52.00
25%     2.75  59.50
50%     4.50  67.00
75%     6.25  75.50
max     8.00  85.00
  • count: how many values.
  • mean: the average. std tells you how spread out the values are.
  • min and max: the smallest and largest.
  • 50% is the median, the middle value.

You can also ask for one number at a time:

Python
print(data["score"].mean())
print(data["hours"].corr(data["score"]).round(3))
Output
67.625
0.998

The correlation is a number between -1 and 1. Near 1 means the two columns rise together. Near 0 means no clear link. Near -1 means one goes up while the other goes down.

Step 4: Draw it

A picture often shows what numbers hide. matplotlib draws charts.

Python
import matplotlib.pyplot as plt

plt.scatter(data["hours"], data["score"])
plt.title("Study time and test score")
plt.xlabel("Hours studied")
plt.ylabel("Score")
plt.show()

You will see dots rising from the lower left to the upper right. That upward slant is the pattern: more study time goes with higher scores.

Careful: Correlation does not prove cause. Maybe students who study more also sleep better or pay more attention. The data shows a link, not the reason.

Step 5: Questions to ask about any dataset

Before you build anything with data, ask:

  1. Where did this data come from? Who collected it and why?
  2. How many examples are there? Eight rows is too few to trust. Real projects use hundreds or thousands.
  3. Who or what is missing? Does it cover everyone it should?
  4. Is anything strange? A score of 500 on a 100-point test is a mistake.

Our eight rows are a teaching example. In the next session we will use them to make a prediction, and you will see why more data would make us more confident.

Exercise

Explore a new table

  1. Create a DataFrame with your own data: for example, 6 people with their age and daily_steps.
  2. Print the first rows with head() and the summary with describe().
  3. Find the average of one column with .mean().
  4. Write one sentence describing a pattern you notice, or say that you don't see one.
Need a hint?

pd.DataFrame({"age": [12, 13, 15], "daily_steps": [8000, 9500, 7000]})

Small project

Study time and scores

Explore a small dataset about how many hours students studied and what score they got. Find out whether more study time goes with higher scores.

  1. Create the DataFrame from the lesson and print head() and describe().
  2. Compute the correlation between hours and score with data["hours"].corr(data["score"]).
  3. Draw a scatter plot with plt.scatter(). Add a title and labels.
  4. Add three of your own rows (invent plausible data) and re-run. Did the pattern change?
  5. Write two sentences about what the data shows and one thing it can't tell us.

Stretch it: Add a third column, sleep_hours, and check whether it also correlates with the score.

Quiz

Check your understanding

Pick one answer for each question. Your score appears right in the page; nothing is sent anywhere.

  1. Question 1What is a DataFrame?
  2. Question 2Which line shows the first five rows of a DataFrame called data?
  3. Question 3What does data["score"].mean() give you?
  4. Question 4A correlation of 0.99 between hours studied and score means…
  5. Question 5Why plot the data before building a model?