Session 2 of 5 · 60 minutes
Looking at Data
Use Python and pandas in Google Colab to load a small dataset, summarize it and spot a pattern.
Goals
By the end of this session you can:
- Open a notebook in Google Colab and run Python cells
- Create a pandas DataFrame and read its rows and columns
- Summarize data with head(), describe() and mean()
- Draw a simple chart and describe the pattern you see
Words to know
- Google Colab
- A free website where you run Python in notebook cells. No installation needed.
- pandas
- A Python library for working with tables of data.
- DataFrame
- A table in pandas, with rows and named columns.
- feature
- A column of facts used to make a prediction.
- correlation
- How strongly two columns move together, from -1 to 1.
The lesson
Step 1: Meet Google Colab
Google Colab (colab.research.google.com) runs Python in your browser. Sign in with a Google account and choose New notebook. A notebook is made of cells. Click in a cell, type code and press Shift + Enter to run it.
print("Hello from Colab")Hello from ColabColab already has the libraries we need, so we can start right away.
Step 2: A table in pandas
We will use a tiny made-up dataset about students: the hours each one studied for a test and the score they got.
import pandas as pd
data = pd.DataFrame({
"hours": [1, 2, 3, 4, 5, 6, 7, 8],
"score": [52, 55, 61, 64, 70, 74, 80, 85],
})
print(data.head()) hours score
0 1 52
1 2 55
2 3 61
3 4 64
4 5 70Each row is one student. Each column is one kind of fact. The numbers on the left (0, 1, 2…) are the row numbers pandas adds for you.
Step 3: Summarizing
describe() calculates a summary for every number column.
print(data.describe().round(2)) hours score
count 8.00 8.00
mean 4.50 67.62
std 2.45 11.72
min 1.00 52.00
25% 2.75 59.50
50% 4.50 67.00
75% 6.25 75.50
max 8.00 85.00count: how many values.mean: the average.stdtells you how spread out the values are.minandmax: the smallest and largest.50%is the median, the middle value.
You can also ask for one number at a time:
print(data["score"].mean())
print(data["hours"].corr(data["score"]).round(3))67.625
0.998The correlation is a number between -1 and 1. Near 1 means the two columns rise together. Near 0 means no clear link. Near -1 means one goes up while the other goes down.
Step 4: Draw it
A picture often shows what numbers hide. matplotlib draws charts.
import matplotlib.pyplot as plt
plt.scatter(data["hours"], data["score"])
plt.title("Study time and test score")
plt.xlabel("Hours studied")
plt.ylabel("Score")
plt.show()You will see dots rising from the lower left to the upper right. That upward slant is the pattern: more study time goes with higher scores.
Careful: Correlation does not prove cause. Maybe students who study more also sleep better or pay more attention. The data shows a link, not the reason.
Step 5: Questions to ask about any dataset
Before you build anything with data, ask:
- Where did this data come from? Who collected it and why?
- How many examples are there? Eight rows is too few to trust. Real projects use hundreds or thousands.
- Who or what is missing? Does it cover everyone it should?
- Is anything strange? A score of 500 on a 100-point test is a mistake.
Our eight rows are a teaching example. In the next session we will use them to make a prediction, and you will see why more data would make us more confident.
Exercise
Explore a new table
- Create a DataFrame with your own data: for example, 6 people with their
ageanddaily_steps. - Print the first rows with
head()and the summary withdescribe(). - Find the average of one column with
.mean(). - Write one sentence describing a pattern you notice, or say that you don't see one.
Need a hint?
pd.DataFrame({"age": [12, 13, 15], "daily_steps": [8000, 9500, 7000]})
Small project
Study time and scores
Explore a small dataset about how many hours students studied and what score they got. Find out whether more study time goes with higher scores.
- Create the DataFrame from the lesson and print
head()anddescribe(). - Compute the correlation between hours and score with
data["hours"].corr(data["score"]). - Draw a scatter plot with
plt.scatter(). Add a title and labels. - Add three of your own rows (invent plausible data) and re-run. Did the pattern change?
- Write two sentences about what the data shows and one thing it can't tell us.
Stretch it: Add a third column, sleep_hours, and check whether it also correlates with the score.
Quiz
Check your understanding
Pick one answer for each question. Your score appears right in the page; nothing is sent anywhere.
Feedback on this lesson
Teachers and students: tell us what worked and what didn't. We read this to improve the curriculum.
Received
Thank you for the feedback.
The Education team reviews lesson feedback to improve the curriculum.
Preview mode: this form is not connected to a destination yet, so nothing was sent. Set its web address in src/data/config.json.